report · 2026-10-07

Can AI automate AI R&D yet?

Why it matters here

Epoch’s initial InnovationEval tests one end-to-end post-training research task. Fable 5 and GPT-5.6 Sol, which showed no detected recall of the reference paper, scored 2% and 15% after run-selection and scope corrections. The same-setting SDPO re-run scored 94% (90% interval 72–100%). Newer paper-aware models scored 40% (Fable 5.1) and 63% (GPT-6 Astra); Fable 5 explicitly given the paper scored 82%. These latter conditions test different information access and are not combined with the main novelty-oriented results. The evaluation exposes limitations in algorithmic research and result reporting, without establishing a general ceiling on AI research.

Read original source ↗

AI-R&D capabilityAI-system improvementEvaluation integrity

What the charts show

Run selection and scope corrections change the reported results

Epoch reports corrected in-scope scores of 2% for Fable 5 and 15% for GPT-5.6 Sol, below their claimed 43% and 71%.

One agent attempt per model. The chart compares claimed scores, scores corrected for selection across similar training runs, and scores additionally corrected for out-of-scope changes. Bars are normalized against SDPO; error bars are estimated 90% intervals for seed/rerun noise. Exact values come from the source’s accessible chart description, not visual digitization. The original chart calls the models uncontaminated; the tracker uses the narrower no detected paper recall, which does not rule out contamination.

Grouped horizontal bars show claimed, run-selection-corrected and in-scope scores for two models.
Unaltered pixel capture of the rendered original chart, including title, legend, axes, caption and Epoch AI CC-BY attribution; captured October 7, 2026. This is a browser capture, not the original downloadable image. Enlarge chart ↗

View original chart and methods ↗
Can AI automate AI R&D yet? · Results: first grouped bar chart

How to interpret this chart
  • One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
  • The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
  • No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
  • Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
  • The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
  • Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
  • This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.
Published values from the chart

Exact labels from Epoch’s accessible chart description. Claimed and partially corrected columns are diagnostic context, not additional scored observations.

Run selection and scope corrections change the reported results
SystemClaimedSelection correctedIn-scope
Claude Fable 543%2%2%
GPT-5.6 Sol71%35%15%

Related benchmarks and indicators

What to keep in mind

  • One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
  • The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
  • No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
  • Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
  • The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
  • Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
  • This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.
  • Fable 5.1’s 40% was largely hyperparameter tuning, whose scope is disputed; Astra’s score partly reflects a remembered SDPO implementation. The supplied-paper 82% condition is replication assistance, not independent discovery.

Related evidence