Tests whether a specialized AI system can predict which of two research ideas will perform better before the experiment.
AI-R&D capability
What this measure tells us
Measures research judgment and idea selection, a component of AI R&D.
Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.
Results by version
1,585 later-paper idea pairs
2025-06-01 · 1585 tasks
This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. Human-verified within-paper pairs from NLP, ML, CV and robotics after the 1 June 2024 GPT-4.1 cutoff; winner from at least three benchmarks.
Metric definitions
Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.
This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. Five NLP topics, six to twelve pairs each; same 45 pairs for predictor and expert comparison.
Metric definitions
Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.
Published results · 45-pair NLP expert-comparison subset
This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. 33 non-tied pairs from 35 unpublished ideas; about half AI-generated and half expert-generated; outcomes implemented by expert researchers.
Metric definitions
Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.
Published results · 33 scored unpublished prompting ideas