benchmark

Research outcome prediction

Tests whether a specialized AI system can predict which of two research ideas will perform better before the experiment.

AI-R&D capability

What this measure tells us

Measures research judgment and idea selection, a component of AI R&D.

  • Prediction is not autonomous idea generation, experimental execution or recursive self-improvement.
  • The human comparator covers a separate 45-pair NLP subset, not the full 1,585-pair test.

Results by version

1,585 later-paper idea pairs

2025-06-01 · 1585 tasks

This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. Human-verified within-paper pairs from NLP, ML, CV and robotics after the 1 June 2024 GPT-4.1 cutoff; winner from at least three benchmarks.

Metric definitions

Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.

Published results · 1,585 later-paper idea pairs
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Fine-tuned GPT-4.1 plus paper retrievalPairwise research-outcome prediction accuracy77.0%2025-06-01research-outcome-full-protocolbenchmark author reported
—Unadapted GPT-4.1, Table 4 baselinePairwise research-outcome prediction accuracy51.4%2025-06-01research-outcome-full-vanilla-protocolbenchmark author reported

45-pair NLP expert-comparison subset

2025-06-01 · 45 tasks

This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. Five NLP topics, six to twelve pairs each; same 45 pairs for predictor and expert comparison.

Metric definitions

Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.

Published results · 45-pair NLP expert-comparison subset
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Fine-tuned GPT-4.1 plus paper retrievalPairwise research-outcome prediction accuracy64.4%2025-06-01research-outcome-nlp-protocolbenchmark author reported

33 scored unpublished prompting ideas

2025-06-01 · 33 tasks

This cohort-specific tracker version isolates reference-point applicability; it is not an upstream revision sequence. 33 non-tied pairs from 35 unpublished ideas; about half AI-generated and half expert-generated; outcomes implemented by expert researchers.

Metric definitions

Pairwise research-outcome prediction accuracy (percent): Accuracy on this specific pair population, not a percentage of AI research automated.

Published results · 33 scored unpublished prompting ideas
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Fine-tuned GPT-4.1 plus paper retrievalPairwise research-outcome prediction accuracy63.6%2025-06-01research-outcome-unpublished-protocolbenchmark author reported

Version lineage

Reference points

human baseline

48.9 · percent

Five-expert majority-vote accuracy over 45 NLP pairs, from 222 valid judgments by 25 researchers.

What this reference means: measured baseline · Where it applies: established

Human result applies only to the NLP subset; not 25 independent full-benchmark runs.

human baseline

60 · percent

Ceiling comparator assembled by selecting the best researcher in each NLP topic on this same test set.

What this reference means: measured baseline · Where it applies: established

Post hoc selection on the evaluation set; not a prospectively chosen expert or equal-effort run.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public taskspartial
public codeunknown
public evaluation serviceunknown

Official sources