Compute efficiency improves on fixed research tasks
Authors fit a 3.0-month compute-multiplier doubling time after December 2025 (95% CI 1.7–5.0 months).
Horizontal axis: model release date, not evaluation date. Vertical axis: multiplier on a logarithmic scale. Points are model estimates with hierarchical-bootstrap 95% intervals; fitted trends use frontier models. Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight. Human reference is 1×. Geometric mean over run multipliers within each task, then over evaluated tasks (seven for Opus 5.5 and Fable 5.1; eight for the other main-cohort models), with per-task bridges between GPT-4, human and Opus 5.0 best-of-run reference envelopes. Adjacent-reference ratios are assumed constant across score levels within each task, but may differ across tasks. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.

View original chart and methods ↗
TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts · Figure 1; Table 2; Sections 2.5 and 3.1; Appendix H
How to interpret this chart
- Measures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.
- Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.
- Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.
- Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.
- Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.
- The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.
- Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.
Published values from the chart
Exact printed Table 2 values for the 20-model Figure 1/4 cohort. Table rounding can differ from the more precise landing-page tooltip. These are source values, not plot estimates. Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight. Main Table 2 estimates are retained; the common-seven-task sensitivity estimates are different results.
| Model / reasoning | Reported multiplier | 95% confidence interval | Scored tasks |
|---|---|---|---|
| Opus 5.5 (max) | 2.30× | 1.15–4.37× | 7 (one refused task excluded) |
| Fable 5.1 (max) | 1.72× | 0.75–3.70× | 7 (one refused task excluded) |
| GPT-6 Astra (max) | 1.60× | 0.77–3.02× | 8 |
| Opus 5.0 (max) | 1.29× | 0.69–2.33× | 8 |
| Opus 4.8 (max) | 0.84× | 0.45–1.78× | 8 |
| GLM 5.3 (max) | 0.59× | 0.33–1.49× | 8 |
| GPT-5.6 Sol (max) | 0.59× | 0.31–1.09× | 8 |
| Opus 4.7 (max) | 0.50× | 0.27–0.95× | 8 |
| GPT-5.5 (xhigh) | 0.47× | 0.24–1.11× | 8 |
| Kimi K3 (reasoning on) | 0.45× | 0.31–0.71× | 8 |
| Opus 4.6 (max) | 0.40× | 0.24–0.85× | 8 |
| GPT-5.2 (high) | 0.28× | 0.16–0.54× | 8 |
| GPT-5.4 (xhigh) | 0.21× | 0.12–0.37× | 8 |
| Opus 4.5 (24k thinking budget) | 0.12× | 0.05–0.24× | 8 |
| GPT-5 (high) | 0.12× | 0.06–0.21× | 8 |
| Kimi K2.5 (max, 96k output) | 0.09× | 0.04–0.19× | 8 |
| o3 (high) | 0.08× | 0.04–0.16× | 8 |
| o1 (high) | 0.07× | 0.03–0.14× | 8 |
| GPT-4o (gpt-4o-2024-05-13) | 0.06× | 0.02–0.23× | 8 |
| GPT-4 (gpt-4-0613) | 0.03× | 0.01–0.08× | 8 |
