Run selection and scope corrections change the reported results
Epoch reports corrected in-scope scores of 2% for Fable 5 and 15% for GPT-5.6 Sol, below their claimed 43% and 71%.
One agent attempt per model. The chart compares claimed scores, scores corrected for selection across similar training runs, and scores additionally corrected for out-of-scope changes. Bars are normalized against SDPO; error bars are estimated 90% intervals for seed/rerun noise. Exact values come from the source’s accessible chart description, not visual digitization. The original chart calls the models uncontaminated; the tracker uses the narrower no detected paper recall, which does not rule out contamination.

View original chart and methods ↗
Can AI automate AI R&D yet? · Results: first grouped bar chart
How to interpret this chart
- One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
- The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
- No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
- Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
- The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
- Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
- This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.
Published values from the chart
Exact labels from Epoch’s accessible chart description. Claimed and partially corrected columns are diagnostic context, not additional scored observations.
| System | Claimed | Selection corrected | In-scope |
|---|---|---|---|
| Claude Fable 5 | 43% | 2% | 2% |
| GPT-5.6 Sol | 71% | 35% | 15% |