Evidence record
InnovationEval
Initial SDPO task · October 7, 2026 · Claude Fable 5 · InnovationEval · 2026-10-07
Correction or update linkedAdded InnovationEval: an end-to-end AI training research task, with corrected results, a human-developed method reference and explicit paper-knowledge limitations.
Reported result · numeric
2%
Estimated seed/rerun confidence interval: 90% CI 0–10%; training noise, one agent attempt
Epoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out.
Epoch-corrected normalized in-scope score · percent
benchmark author reported · extraction review: agent checked
Sources
Published 2026-10-07
Evaluation setup
| split | Visible short-answer test sets; transductive coding evaluation, No detected paper recall describes the information-access condition, not the evaluation split. |
|---|---|
| task count | 1 |
| task snapshot | One ML research task: one method applied to five short-answer datasets and one coding dataset (six datasets are not six independent research tasks). Chemistry, physics, biology, materials and tool-use test sets contain 210, 80, 50, 94 and 68 items; coding uses 131 transductive problems. |
| attempts per task | 1 |
| run count | 1, One main research attempt per model, containing many training experiments; this is not a count of training runs. |
| selection rule | One main agent submission per model. Epoch averages similar training runs to correct best-of-run selection and estimates in-scope effects from experiment logs. |
| aggregate method | Two equally weighted areas: short-answer science/tool use and coding. Short-answer combines pooled five-dataset improvement at one hour with the five-hour parity score 1 + z_score/6 clipped between 0 and 1 (zero at a −3.6 percentage-point mean gap, one at parity); coding combines final and average performance against 20,480 training generations. Epoch rescales raw composite x on a 0–100 scale as 100 × (x − 25) / 75, then reports corrections for best-of-run selection and out-of-scope changes. Coding scope correction compares similar wall-clock times. |
| token budget | 10000000000 |
| hardware | Modal GPU jobs, default H200 141 GB; maximum 50 concurrent GPUs. Qwen3-8B target model with thinking off. |
| training budget | 3,000 GPU-hours per agent attempt; Fable 5 used 46%, GPT-5.6 Sol used the full budget. Reported GPU expenditure: approximately $6,700 and $14,000 respectively. |
| inference budget | 10 billion input + output + reasoning tokens; consumption differs between models. Target evaluation avg@16 for short-answer and avg@4 for coding. Reported usage: Fable 5 1.8% ($610), GPT-5.6 Sol 24% ($2,100). |
| monetary cost | not reported, All-inclusive research cost is not reported. Reported components: Claude Fable 5 about $6,700 GPU usage and $610 inference tokens; GPT-5.6 Sol about $14,000 GPU usage and $2,100 inference tokens. These are separate consumed-resource components, not matched total research costs. |
| tool access | Inspect ReAct with bash, text editor and tools to submit and monitor GPU jobs; starting verl codebase with SDPO removed. |
| internet access | false |
| filtering | Algorithmic changes restricted to loss, updates and use of current-batch rollouts; no extra training data or task-specific methods. |
| scaffold | Inspect ReAct |
| human intervention | Humans define task, constraints, baseline and references, then review submissions, timing logs and transcripts with LLM assistance; no human-guided agent research reported. |
| contamination concerns | No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication. Newer paper-aware models and a supplied-paper replication attempt are retained separately as contextual figure evidence. |
| comparability caveats | One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.; The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.; No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.; Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.; The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.; Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.; This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.; The prompt permits unseen-dataset regrading; this preliminary report does not demonstrate that such regrading occurred. |
Not reported: wall clock budget, evaluator version, task exclusions.
Comparability
Not compared with other results.
Limitations
- One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
- The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
- No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
- Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
- The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
- Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
- This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.