Initial SDPO task · October 7, 2026
2026-10-07 · 1 tasksTwo equally weighted areas: short-answer science/tool use and coding. Short-answer combines pooled five-dataset improvement at one hour with the five-hour parity score 1 + z_score/6 clipped between 0 and 1 (zero at a −3.6 percentage-point mean gap, one at parity); coding combines final and average performance against 20,480 training generations. Epoch rescales raw composite x on a 0–100 scale as 100 × (x − 25) / 75, then reports corrections for best-of-run selection and out-of-scope changes. Coding scope correction compares similar wall-clock times.
Metric definitions
Epoch-corrected normalized in-scope score (percent): Normalized to GRPO at 0% and the SDPO reference at 100%. Negative scores are possible. The reference re-run achieved 94% (90% CI 72–100%). A score is not a direct measure of novelty.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| — | Claude Fable 5 · InnovationEval | Epoch-corrected normalized in-scope score | 2%Estimated seed/rerun confidence interval: 90% CI 0–10%; training noise, one agent attemptEpoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out. | 2026-10-07 | innovationeval-no-detected-recall | benchmark author reported | |
| — | GPT-5.6 Sol · InnovationEval | Epoch-corrected normalized in-scope score | 15%Estimated seed/rerun confidence interval: 90% CI 2–28%; training noise, one agent attemptEpoch adjusted these scores for run selection and method-scope judgments. One main research attempt was evaluated per model; no paper recall was detected, but contamination is not ruled out. | 2026-10-07 | innovationeval-no-detected-recall | benchmark author reported |
No results match these filters for this version.