preprint · 2025-06-01
Predicting Empirical AI Research Outcomes with Language Models
Why it matters here
A specialized fine-tuned GPT-4.1 and paper-retrieval system predicted which of two ideas performed better: 77.0% on 1,585 later-paper pairs. On a separate 45-pair NLP subset it scored 64.4% versus 48.9% five-expert majority vote; 33 unpublished pairs yielded 63.6%. This measures prediction before experimentation, not autonomous end-to-end research.
AI-R&D capability
Related benchmarks and indicators
What to keep in mind
- The expert comparison is limited to 45 NLP pairs and includes 222 valid researcher judgments, not the full test set.
- The post hoc best-expert-per-topic reference of 60.0% was selected on the same 45 pairs.
- The paper reports 51.9% for unadapted GPT-4.1 in Figure 3/prose and 51.4% in Table 4; baseline configurations cannot be safely equated.
- Research ideas and empirical outcomes were supplied by papers or human implementers; the predictor did not run the experiments.