Evidence record
Research outcome prediction
1,585 later-paper idea pairs · Fine-tuned GPT-4.1 plus paper retrieval · 2025-06-01
Reported result · numeric
77.0%
Pairwise research-outcome prediction accuracy · percent
benchmark author reported · extraction review: agent checked
Sources
Published 2025-06-01
Evaluation setup
| split | Published-paper test after GPT-4.1 cutoff |
|---|---|
| task count | 1585 |
| task snapshot | Source paper arXiv:2506.00794v1; GPT-4.1 pre/post 1 June 2024 cutoff as applicable |
| attempts per task | 2 |
| selection rule | Predict both idea orders; count inconsistent answers incorrect |
| aggregate method | Percent correct on eligible idea pairs |
| tool access | Paper retrieval agent, before-cutoff literature only |
| filtering | Same-paper idea pairs; at least three benchmarks; ties removed |
| scaffold | Fine-tuned GPT-4.1 on 6,000 historical pair labels plus iterative paper retrieval |
| contamination concerns | Retrieval limited to papers dated before 1 June 2024 |
Not reported: run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, internet access, evaluator version, human intervention, task exclusions, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Specialized fine-tuning and retrieval; not an off-the-shelf GPT-4.1 score.
- No matched full-test human baseline.
- Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
- The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.