October 2026 paper v1 · 20-model main cohort
2026-10-05 · 8 tasksPrivate eight-task suite; six seeds per model per task. Records copy Table 2 at two-decimal precision for the 20 models plotted in Figures 1 and 4. Separate effort/Super-Think ablations and GPT-3.5/Llama supplementary rows are not included in this cohort. Model release dates on the original figures are not evaluation dates. Retain both metrics and uncertainty. Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight. Main Table 2 values are retained; the separate common-seven-task sensitivity is not substituted.
Metric definitions
Experimental compute multiplier (×): Effective serial experimental compute in human-reference units (1×), using best-of-run envelopes, the weaker endpoint and chained references. An envelope may combine different human runs. Excludes Researcher inference tokens. Not a general intelligence or RSI score.
Normalized final performance multiplier (×): Final test score rescaled so naive baseline=0 and expert baseline=1, regardless of experimental compute consumed. Different from the compute multiplier; not percent accuracy.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| — | Opus 5.5 (max) / TasteVal | Experimental compute multiplier | 2.30×Hierarchical bootstrap: 95% CI 1.15–4.37×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero. | 2026-10-05 | tasteval-v1-compute-multiplier-seven-task-protocol | benchmark author reported | |
| — | Opus 5.5 (max) / TasteVal | Normalized final performance multiplier | 1.14×Hierarchical bootstrap: 95% CI 1.01–1.31×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero. | 2026-10-05 | tasteval-v1-performance-multiplier-seven-task-protocol | benchmark author reported | |
| — | Fable 5.1 (max) / TasteVal | Experimental compute multiplier | 1.72×Hierarchical bootstrap: 95% CI 0.75–3.70×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero. | 2026-10-05 | tasteval-v1-compute-multiplier-seven-task-protocol | benchmark author reported | |
| — | Fable 5.1 (max) / TasteVal | Normalized final performance multiplier | 1.09×Hierarchical bootstrap: 95% CI 0.91–1.29×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero. | 2026-10-05 | tasteval-v1-performance-multiplier-seven-task-protocol | benchmark author reported | |
| — | GPT-6 Astra (max) / TasteVal | Experimental compute multiplier | 1.60×Hierarchical bootstrap: 95% CI 0.77–3.02× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-6 Astra (max) / TasteVal | Normalized final performance multiplier | 1.13×Hierarchical bootstrap: 95% CI 0.94–1.32× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Opus 5.0 (max) / TasteVal | Experimental compute multiplier | 1.29×Hierarchical bootstrap: 95% CI 0.69–2.33× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Opus 5.0 (max) / TasteVal | Normalized final performance multiplier | 1.11×Hierarchical bootstrap: 95% CI 0.91–1.35× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Opus 4.8 (max) / TasteVal | Experimental compute multiplier | 0.84×Hierarchical bootstrap: 95% CI 0.45–1.78× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Opus 4.8 (max) / TasteVal | Normalized final performance multiplier | 1.01×Hierarchical bootstrap: 95% CI 0.80–1.37× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GLM 5.3 (max) / TasteVal | Experimental compute multiplier | 0.59×Hierarchical bootstrap: 95% CI 0.33–1.49× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GLM 5.3 (max) / TasteVal | Normalized final performance multiplier | 0.89×Hierarchical bootstrap: 95% CI 0.58–1.26× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-5.6 Sol (max) / TasteVal | Experimental compute multiplier | 0.59×Hierarchical bootstrap: 95% CI 0.31–1.09× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-5.6 Sol (max) / TasteVal | Normalized final performance multiplier | 0.92×Hierarchical bootstrap: 95% CI 0.71–1.15× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Opus 4.7 (max) / TasteVal | Experimental compute multiplier | 0.50×Hierarchical bootstrap: 95% CI 0.27–0.95× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Opus 4.7 (max) / TasteVal | Normalized final performance multiplier | 0.75×Hierarchical bootstrap: 95% CI 0.42–1.09× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-5.5 (xhigh) / TasteVal | Experimental compute multiplier | 0.47×Hierarchical bootstrap: 95% CI 0.24–1.11× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-5.5 (xhigh) / TasteVal | Normalized final performance multiplier | 0.84×Hierarchical bootstrap: 95% CI 0.55–1.13× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Kimi K3 (reasoning on) / TasteVal | Experimental compute multiplier | 0.45×Hierarchical bootstrap: 95% CI 0.31–0.71× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Kimi K3 (reasoning on) / TasteVal | Normalized final performance multiplier | 0.89×Hierarchical bootstrap: 95% CI 0.71–1.12× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Opus 4.6 (max) / TasteVal | Experimental compute multiplier | 0.40×Hierarchical bootstrap: 95% CI 0.24–0.85× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Opus 4.6 (max) / TasteVal | Normalized final performance multiplier | 0.78×Hierarchical bootstrap: 95% CI 0.49–1.14× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-5.2 (high) / TasteVal | Experimental compute multiplier | 0.28×Hierarchical bootstrap: 95% CI 0.16–0.54× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-5.2 (high) / TasteVal | Normalized final performance multiplier | 0.70×Hierarchical bootstrap: 95% CI 0.39–0.98× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-5.4 (xhigh) / TasteVal | Experimental compute multiplier | 0.21×Hierarchical bootstrap: 95% CI 0.12–0.37× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-5.4 (xhigh) / TasteVal | Normalized final performance multiplier | 0.60×Hierarchical bootstrap: 95% CI 0.36–0.84× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Opus 4.5 (24k thinking budget) / TasteVal | Experimental compute multiplier | 0.12×Hierarchical bootstrap: 95% CI 0.05–0.24× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Opus 4.5 (24k thinking budget) / TasteVal | Normalized final performance multiplier | 0.37×Hierarchical bootstrap: 95% CI 0.11–0.79× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-5 (high) / TasteVal | Experimental compute multiplier | 0.12×Hierarchical bootstrap: 95% CI 0.06–0.21× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-5 (high) / TasteVal | Normalized final performance multiplier | 0.53×Hierarchical bootstrap: 95% CI 0.22–0.73× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | Kimi K2.5 (max, 96k output) / TasteVal | Experimental compute multiplier | 0.09×Hierarchical bootstrap: 95% CI 0.04–0.19× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | Kimi K2.5 (max, 96k output) / TasteVal | Normalized final performance multiplier | 0.23×Hierarchical bootstrap: 95% CI 0.06–0.69× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | o3 (high) / TasteVal | Experimental compute multiplier | 0.08×Hierarchical bootstrap: 95% CI 0.04–0.16× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | o3 (high) / TasteVal | Normalized final performance multiplier | 0.43×Hierarchical bootstrap: 95% CI 0.18–0.61× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | o1 (high) / TasteVal | Experimental compute multiplier | 0.07×Hierarchical bootstrap: 95% CI 0.03–0.14× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | o1 (high) / TasteVal | Normalized final performance multiplier | 0.22×Hierarchical bootstrap: 95% CI 0.06–0.48× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-4o (gpt-4o-2024-05-13) / TasteVal | Experimental compute multiplier | 0.06×Hierarchical bootstrap: 95% CI 0.02–0.23× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-4o (gpt-4o-2024-05-13) / TasteVal | Normalized final performance multiplier | 0.29×Hierarchical bootstrap: 95% CI 0.09–0.59× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported | |
| — | GPT-4 (gpt-4-0613) / TasteVal | Experimental compute multiplier | 0.03×Hierarchical bootstrap: 95% CI 0.01–0.08× | 2026-10-05 | tasteval-v1-compute-multiplier-protocol | benchmark author reported | |
| — | GPT-4 (gpt-4-0613) / TasteVal | Normalized final performance multiplier | 0.22×Hierarchical bootstrap: 95% CI 0.06–0.55× | 2026-10-05 | tasteval-v1-performance-multiplier-protocol | benchmark author reported |
No results match these filters for this version.