Evidence record

Research outcome prediction

33 scored unpublished prompting ideas · Fine-tuned GPT-4.1 plus paper retrieval · 2025-06-01

Reported result · numeric

63.6%

Pairwise research-outcome prediction accuracy · percent

benchmark author reported · extraction review: agent checked

Sources

Published 2025-06-01

Evaluation setup
split33 non-tied unpublished idea/baseline pairs
task count33
task snapshotSource paper arXiv:2506.00794v1; GPT-4.1 pre/post 1 June 2024 cutoff as applicable
attempts per task2
selection rulePredict both idea orders; count inconsistent answers incorrect
aggregate methodPercent correct on eligible idea pairs
tool accessPaper retrieval agent, before-cutoff literature only
filteringSame-paper idea pairs; at least three benchmarks; ties removed
scaffoldFine-tuned GPT-4.1 on 6,000 historical pair labels plus iterative paper retrieval
human interventionExpert researchers implemented the proposed ideas, averaging 103.4 hours per original idea
contamination concernsRetrieval limited to papers dated before 1 June 2024

Not reported: run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, internet access, evaluator version, task exclusions, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • 33 scored pairs after removing ties from 35 initial ideas; small prompting-method transfer test.
  • Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
  • The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.