Evidence record
Research outcome prediction
33 scored unpublished prompting ideas · Fine-tuned GPT-4.1 plus paper retrieval · 2025-06-01
Reported result · numeric
63.6%
Pairwise research-outcome prediction accuracy · percent
benchmark author reported · extraction review: agent checked
Sources
Published 2025-06-01
Evaluation setup
| split | 33 non-tied unpublished idea/baseline pairs |
|---|---|
| task count | 33 |
| task snapshot | Source paper arXiv:2506.00794v1; GPT-4.1 pre/post 1 June 2024 cutoff as applicable |
| attempts per task | 2 |
| selection rule | Predict both idea orders; count inconsistent answers incorrect |
| aggregate method | Percent correct on eligible idea pairs |
| tool access | Paper retrieval agent, before-cutoff literature only |
| filtering | Same-paper idea pairs; at least three benchmarks; ties removed |
| scaffold | Fine-tuned GPT-4.1 on 6,000 historical pair labels plus iterative paper retrieval |
| human intervention | Expert researchers implemented the proposed ideas, averaging 103.4 hours per original idea |
| contamination concerns | Retrieval limited to papers dated before 1 June 2024 |
Not reported: run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, internet access, evaluator version, task exclusions, comparability caveats.
Comparability
Not compared with other results.
Limitations
- 33 scored pairs after removing ties from 35 initial ideas; small prompting-method transfer test.
- Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
- The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.