benchmark

InnovationEval

Tests whether an AI agent can develop a post-training method approaching a recent human innovation, from ideas through implementation and experiments.

AI-R&D capabilityAI-system improvementEvaluation integrity

What this measure tells us

Directly tests an important ingredient of AI self-improvement: end-to-end algorithmic research on AI training. The first task targets SDPO-like gains over GRPO while post-training Qwen3-8B.

  • One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.
  • The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.
  • No detected paper recall is not proof of uncontaminated training: GPT-5.6 Sol’s cutoff overlaps the paper’s January 28 publication.
  • Human reviewers assessed submissions, logs and transcripts with LLM assistance. Scope corrections were estimated from existing experiments; footnote 15 says this preliminary report did not require fresh re-grading with scope violations ablated.
  • The coding task is transductive: training and evaluation use the same 131 problems with different unit-test access. Test splits were visible for selection.
  • Coding scope correction uses comparable wall-clock time, although the task prompt specifies a generation-based axis; Epoch acknowledges this judgment in footnote 20.
  • This is normalized task performance, not percent of all AI research automated, a direct novelty grade, a matched human research success rate, or evidence of recursive acceleration.

Results by version

Initial SDPO task · October 7, 2026

2026-10-07 · 1 tasks

Two equally weighted areas: short-answer science/tool use and coding. Short-answer combines pooled five-dataset improvement at one hour with the five-hour parity score 1 + z_score/6 clipped between 0 and 1 (zero at a −3.6 percentage-point mean gap, one at parity); coding combines final and average performance against 20,480 training generations. Epoch rescales raw composite x on a 0–100 scale as 100 × (x − 25) / 75, then reports corrections for best-of-run selection and out-of-scope changes. Coding scope correction compares similar wall-clock times.

Metric definitions

Epoch-corrected normalized in-scope score (percent): Normalized to GRPO at 0% and the SDPO reference at 100%. Negative scores are possible. The reference re-run achieved 94% (90% CI 72–100%). A score is not a direct measure of novelty.

Published results · Initial SDPO task · October 7, 2026
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Claude Fable 5 · InnovationEvalEpoch-corrected normalized in-scope score2%Estimated seed/rerun confidence interval: 90% CI 0–10%; training noise, one agent attemptEpoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out.2026-10-07innovationeval-no-detected-recallbenchmark author reported
—GPT-5.6 Sol · InnovationEvalEpoch-corrected normalized in-scope score15%Estimated seed/rerun confidence interval: 90% CI 2–28%; training noise, one agent attemptEpoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out.2026-10-07innovationeval-no-detected-recallbenchmark author reported

Version lineage

Reference points

normalization anchor

100 · percent

SDPO normalization target, not a measured human research success rate or a sufficient condition for RSI.

What this reference means: unknown · Where it applies: established

One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.

The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.

other

94 · percent

SDPO method re-run: 94% (90% interval 72–100%). Epoch re-ran the original method in this setup; this measures training-method performance, not human researchers operating under the agent budget.

What this reference means: measured baseline · Where it applies: established

One research task based on one human paper, with one main agent attempt per model; this does not establish general research reliability.

The 90% confidence intervals concern seed/rerun noise in training the submitted method, not variability across independent agent attempts.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public taskspartial
public codeunknown
public evaluation serviceunknown

Evidence in source charts

Related benchmarks

  • TasteVal, a separate benchmark version with its own results.
  • PostTrainBench, a separate benchmark version with its own results.

Official sources