Evidence record

Research outcome prediction

1,585 later-paper idea pairs · Fine-tuned GPT-4.1 plus paper retrieval · 2025-06-01

Reported result · numeric

77.0%

Pairwise research-outcome prediction accuracy · percent

benchmark author reported · extraction review: agent checked

Sources

Published 2025-06-01

Evaluation setup
splitPublished-paper test after GPT-4.1 cutoff
task count1585
task snapshotSource paper arXiv:2506.00794v1; GPT-4.1 pre/post 1 June 2024 cutoff as applicable
attempts per task2
selection rulePredict both idea orders; count inconsistent answers incorrect
aggregate methodPercent correct on eligible idea pairs
tool accessPaper retrieval agent, before-cutoff literature only
filteringSame-paper idea pairs; at least three benchmarks; ties removed
scaffoldFine-tuned GPT-4.1 on 6,000 historical pair labels plus iterative paper retrieval
contamination concernsRetrieval limited to papers dated before 1 June 2024

Not reported: run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, internet access, evaluator version, human intervention, task exclusions, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • Specialized fine-tuning and retrieval; not an off-the-shelf GPT-4.1 score.
  • No matched full-test human baseline.
  • Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
  • The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.