preprint · 2025-06-01

Predicting Empirical AI Research Outcomes with Language Models

Why it matters here

A specialized fine-tuned GPT-4.1 and paper-retrieval system predicted which of two ideas performed better: 77.0% on 1,585 later-paper pairs. On a separate 45-pair NLP subset it scored 64.4% versus 48.9% five-expert majority vote; 33 unpublished pairs yielded 63.6%. This measures prediction before experimentation, not autonomous end-to-end research.

Read original source ↗

AI-R&D capability

Related benchmarks and indicators

What to keep in mind

  • The expert comparison is limited to 45 NLP pairs and includes 222 valid researcher judgments, not the full test set.
  • The post hoc best-expert-per-topic reference of 60.0% was selected on the same 45 pairs.
  • The paper reports 51.9% for unadapted GPT-4.1 in Figure 3/prose and 51.4% in Table 4; baseline configurations cannot be safely equated.
  • Research ideas and empirical outcomes were supplied by papers or human implementers; the predictor did not run the experiments.

Related evidence