Evidence record
TasteVal
October 2026 paper v1 · 20-model main cohort · GPT-5.2 (high) / TasteVal · 2026-10-05
Reported result · numeric
0.28×
Hierarchical bootstrap: 95% CI 0.16–0.54×
Experimental compute multiplier · ×
benchmark author reported · extraction review: agent checked
Sources
Published 2026-10-05
Evaluation setup
| split | Train and validation visible to Researcher; test/scoring visible to Coder only. Final result uses official-submission test score. |
|---|---|
| task count | 8 |
| task snapshot | TasteVal arXiv:2610.06824v1 eight-scored-task main cohort, October 2026. Six seeds per task imply 48 included task-seed results; this does not establish the total number of attempted sessions. |
| attempts per task | 6 |
| run count | 48 |
| selection rule | Researcher-designated official submission; otherwise best validation submission. Confirmed reward-hacking submissions discarded; fallback to best clean validation submission, then naive floor if none. |
| aggregate method | Geometric mean over run multipliers within each task, then over evaluated tasks (seven for Opus 5.5 and Fable 5.1; eight for the other main-cohort models), with per-task bridges between GPT-4, human and Opus 5.0 best-of-run reference envelopes. Adjacent-reference ratios are assumed constant across score levels within each task, but may differ across tasks. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed. |
| wall clock budget | 120 hours per attempt, or 40 H100 GPU-busy hours, whichever is exhausted first. |
| hardware | One H100 per attempt; experiments run serially. |
| training budget | 40 H100 GPU-busy hours, including experiments and scoring. |
| inference budget | Researcher token use is outside the GPU budget, subject to the wall-clock limit; effort varies by model as labeled. |
| tool access | Researcher plans, reads task files/reports and metrics, uses internet and manages experiments; fixed Opus 4.8 Coder implements and executes. Two monitors and a liaison mediate exchanges. |
| internet access | true |
| filtering | Automated audit followed by manual confirmation of flagged submissions; no independent tracker audit of all trajectories. |
| scaffold | TasteVal ReAct; GPT-4 and GPT-4o use ReAct-Shim. Opus 4.8 Coder fixed across humans and models. |
| evaluator version | Private TasteVal v1 scoring scripts; reported hierarchical bootstrap with 1,000 task-then-run resamples. |
| human intervention | Humans authored tasks, prompts and rules, supplied expert baselines, and confirmed flagged violations. Recruited experts had 4–16 hours of preliminary literature review and no AI help with experimental reasoning. |
| task exclusions | This cohort omits separate effort/Super-Think ablations and the two oldest supplemental models. Opus 5.5 and Fable 5.1 use the separate seven-scored-task protocol. |
| contamination concerns | Authors created and withheld tasks; this design does not independently prove absence of all leakage. |
| comparability caveats | Measures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.; Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.; Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.; Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.; Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.; The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.; Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.; Two oldest models use a different scaffold. Confidence intervals do not establish pairwise model differences. |
Not reported: token budget, monetary cost.
Comparability
Not compared with other results.
Limitations
- Measures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.
- Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.
- Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.
- Source-reported confidence intervals and rounded Table 2 values; overlapping intervals do not prove pairwise differences.
- Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.
- Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.
- The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.
- Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.