Evidence record

InnovationEval

Initial SDPO task · October 7, 2026 · GPT-5.6 Sol · InnovationEval · 2026-10-07

Correction or update linkedAdded InnovationEval: an end-to-end AI training research task, with corrected results, a human-developed method reference and explicit paper-knowledge limitations.

Reported result · numeric

15%

Estimated seed/rerun confidence interval: 90% CI 2–28%; training noise, one agent attempt

Epoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out.

Epoch-corrected normalized in-scope score · percent

benchmark author reported · extraction review: agent checked

Sources

Published 2026-10-07

Evaluation setup
splitVisible short-answer test sets; transductive coding evaluation, No detected paper recall describes the information-access condition, not the evaluation split.
task count1
task snapshotOne ML research task: one method applied to five short-answer datasets and one coding dataset (six datasets are not six independent research tasks). Chemistry, physics, biology, materials and tool-use test sets contain 210, 80, 50, 94 and 68 items; coding uses 131 transductive problems.
attempts per task1
run count1, One main research attempt per model, containing many training experiments; this is not a count of training runs.
selection ruleOne main agent submission per model. Epoch averages similar training runs to correct best-of-run selection and estimates in-scope effects from experiment logs.
aggregate methodTwo equally weighted areas: short-answer science/tool use and coding. Short-answer combines pooled five-dataset improvement at one hour with the five-hour parity score 1 + z_score/6 clipped between 0 and 1 (zero at a −3.6 percentage-point mean gap, one at parity); coding combines final and average performance against 20,480 training generations. Epoch rescales raw composite x on a 0–100 scale as 100 × (x − 25) / 75, then reports corrections for best-of-run selection and out-of-scope changes. Coding scope correction compares similar wall-clock times.
token budget10000000000
hardwareModal GPU jobs, default H200 141 GB; maximum 50 concurrent GPUs. Qwen3-8B target model with thinking off.
training budget3,000 GPU-hours per agent attempt; Fable 5 used 46%, GPT-5.6 Sol used the full budget. Reported GPU expenditure: approximately $6,700 and $14,000 respectively.
inference budget10 billion input + output + reasoning tokens; consumption differs between models. Target evaluation avg@16 for short-answer and avg@4 for coding. Reported usage: Fable 5 1.8% ($610), GPT-5.6 Sol 24% ($2,100).
monetary costnot reported, All-inclusive research cost is not reported. Reported components: Claude Fable 5 about $6,700 GPU usage and $610 inference tokens; GPT-5.6 Sol about $14,000 GPU usage and $2,100 inference tokens. These are separate consumed-resource components, not matched total research costs.
tool accessInspect ReAct with bash, text editor and tools to submit and monitor GPU jobs; starting verl codebase with SDPO removed.
internet accessfalse
filteringAlgorithmic changes restricted to loss, updates and use of current-batch rollouts; no extra training data or task-specific methods.
scaffoldInspect ReAct
human interventionHumans define task, constraints, baseline and references, then review submissions, timing logs and transcripts with LLM assistance; no human-guided agent research reported.
contamination concernsNo detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication. Newer paper-aware models and a supplied-paper replication attempt are retained separately as contextual figure evidence.
comparability caveatsOne research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.; The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.; No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.; Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.; The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.; Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.; This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.; The prompt permits unseen-dataset regrading; this preliminary report does not demonstrate that such regrading occurred.

Not reported: wall clock budget, evaluator version, task exclusions.

Comparability

Not compared with other results.

Limitations

  • One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
  • The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
  • No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
  • Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
  • The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
  • Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
  • This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.