Evidence record

TasteVal

October 2026 paper v1 · 20-model main cohort · GPT-5.5 (xhigh) / TasteVal · 2026-10-05

Reported result · numeric

0.84×

Hierarchical bootstrap: 95% CI 0.55–1.13×

Normalized final performance multiplier · ×

benchmark author reported · extraction review: agent checked

Sources

Published 2026-10-05

Evaluation setup
splitTrain and validation visible to Researcher; test/scoring visible to Coder only. Final result uses official-submission test score.
task count8
task snapshotTasteVal arXiv:2610.06824v1 eight-scored-task main cohort, October 2026. Six seeds per task imply 48 included task-seed results; this does not establish the total number of attempted sessions.
attempts per task6
run count48
selection ruleResearcher-designated official submission; otherwise best validation submission. Confirmed reward-hacking submissions discarded; fallback to best clean validation submission, then naive floor if none.
aggregate methodArithmetic mean of nonnegative normalized final scores across seeds per task, then geometric mean across evaluated tasks; task floor 0.01. Opus 5.5 and Fable 5.1 omit one refused task; this is missing data, not a zero score.
wall clock budget120 hours per attempt, or 40 H100 GPU-busy hours, whichever is exhausted first.
hardwareOne H100 per attempt; experiments run serially.
training budget40 H100 GPU-busy hours, including experiments and scoring.
inference budgetResearcher token use is outside the GPU budget, subject to the wall-clock limit; effort varies by model as labeled.
tool accessResearcher plans, reads task files/reports and metrics, uses internet and manages experiments; fixed Opus 4.8 Coder implements and executes. Two monitors and a liaison mediate exchanges.
internet accesstrue
filteringAutomated audit followed by manual confirmation of flagged submissions; no independent tracker audit of all trajectories.
scaffoldTasteVal ReAct; GPT-4 and GPT-4o use ReAct-Shim. Opus 4.8 Coder fixed across humans and models.
evaluator versionPrivate TasteVal v1 scoring scripts; reported hierarchical bootstrap with 1,000 task-then-run resamples.
human interventionHumans authored tasks, prompts and rules, supplied expert baselines, and confirmed flagged violations. Recruited experts had 4–16 hours of preliminary literature review and no AI help with experimental reasoning.
task exclusionsThis cohort omits separate effort/Super-Think ablations and the two oldest supplemental models. Opus 5.5 and Fable 5.1 use the separate seven-scored-task protocol.
contamination concernsAuthors created and withheld tasks; this design does not independently prove absence of all leakage.
comparability caveatsMeasures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.; Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.; Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.; Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.; Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.; The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.; Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.; Two oldest models use a different scaffold. Confidence intervals do not establish pairwise model differences.

Not reported: token budget, monetary cost.

Comparability

Not compared with other results.

Limitations

  • Measures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.
  • Human reference is the best attempt per task from 24 recruited experts, at least two per task, not the best researchers worldwide. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.
  • Normalized score multiplier, not compute savings. Runs below naive baseline count as zero before averaging; per-task floor 0.01.
  • Source-reported confidence intervals and rounded Table 2 values; overlapping intervals do not prove pairwise differences.
  • Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.
  • Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.
  • Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.
  • Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.
  • The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.
  • Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.