benchmark

TasteVal

Measures how efficiently models choose experiments and interpret results on eight fixed AI R&D problems, with a shared coding agent and human expert reference.

AI-R&D capabilityAI-system improvementEvaluation integrity

What this measure tells us

Tests one component of AI R&D: experimental planning and feedback-driven optimization under a fixed budget.

  • Measures experimental planning on fixed problems with fast feedback, not choosing worthwhile problems or general scientific taste.
  • Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.
  • Eight private tasks on one H100; transfer to frontier-scale, parallel or open-ended research is unestablished. Tasks are withheld, limiting independent reproduction.
  • Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.
  • Compute multipliers use chained best-of-run reference envelopes and geometric aggregation; they are not a direct model-versus-human runtime ratio on every task. Bootstrap bridge ratios are re-estimated while reference paths, bridge membership and envelopes stay fixed.
  • The 3.0-month fitted doubling time applies to compute efficiency after December 2025; normalized final performance doubles every 14.6 months without a significant trend break. Neither trend establishes recursive acceleration.
  • Authors found no Invented submissions in their labeled sample; progress mainly involved tuning, composition and moderate adaptation. This is not evidence that models cannot invent in other settings.

Results by version

October 2026 paper v1 · 20-model main cohort

2026-10-05 · 8 tasks

Private eight-task suite; six seeds per model per task. Records copy Table 2 at two-decimal precision for the 20 models plotted in Figures 1 and 4. Separate effort/Super-Think ablations and GPT-3.5/Llama supplementary rows are not included in this cohort. Model release dates on the original figures are not evaluation dates. Retain both metrics and uncertainty. Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight. Main Table 2 values are retained; the separate common-seven-task sensitivity is not substituted.

Metric definitions

Experimental compute multiplier (×): Effective serial experimental compute in human-reference units (1×), using best-of-run envelopes, the weaker endpoint and chained references. An envelope may combine different human runs. Excludes Researcher inference tokens. Not a general intelligence or RSI score.

Normalized final performance multiplier (×): Final test score rescaled so naive baseline=0 and expert baseline=1, regardless of experimental compute consumed. Different from the compute multiplier; not percent accuracy.

Published results · October 2026 paper v1 · 20-model main cohort
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Opus 5.5 (max) / TasteValExperimental compute multiplier2.30×Hierarchical bootstrap: 95% CI 1.15–4.37×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero.2026-10-05tasteval-v1-compute-multiplier-seven-task-protocolbenchmark author reported
—Opus 5.5 (max) / TasteValNormalized final performance multiplier1.14×Hierarchical bootstrap: 95% CI 1.01–1.31×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero.2026-10-05tasteval-v1-performance-multiplier-seven-task-protocolbenchmark author reported
—Fable 5.1 (max) / TasteValExperimental compute multiplier1.72×Hierarchical bootstrap: 95% CI 0.75–3.70×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero.2026-10-05tasteval-v1-compute-multiplier-seven-task-protocolbenchmark author reported
—Fable 5.1 (max) / TasteValNormalized final performance multiplier1.09×Hierarchical bootstrap: 95% CI 0.91–1.29×Scored on seven tasks; one refused task was excluded as missing data, not failure or zero.2026-10-05tasteval-v1-performance-multiplier-seven-task-protocolbenchmark author reported
—GPT-6 Astra (max) / TasteValExperimental compute multiplier1.60×Hierarchical bootstrap: 95% CI 0.77–3.02×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-6 Astra (max) / TasteValNormalized final performance multiplier1.13×Hierarchical bootstrap: 95% CI 0.94–1.32×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Opus 5.0 (max) / TasteValExperimental compute multiplier1.29×Hierarchical bootstrap: 95% CI 0.69–2.33×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Opus 5.0 (max) / TasteValNormalized final performance multiplier1.11×Hierarchical bootstrap: 95% CI 0.91–1.35×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Opus 4.8 (max) / TasteValExperimental compute multiplier0.84×Hierarchical bootstrap: 95% CI 0.45–1.78×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Opus 4.8 (max) / TasteValNormalized final performance multiplier1.01×Hierarchical bootstrap: 95% CI 0.80–1.37×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GLM 5.3 (max) / TasteValExperimental compute multiplier0.59×Hierarchical bootstrap: 95% CI 0.33–1.49×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GLM 5.3 (max) / TasteValNormalized final performance multiplier0.89×Hierarchical bootstrap: 95% CI 0.58–1.26×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-5.6 Sol (max) / TasteValExperimental compute multiplier0.59×Hierarchical bootstrap: 95% CI 0.31–1.09×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-5.6 Sol (max) / TasteValNormalized final performance multiplier0.92×Hierarchical bootstrap: 95% CI 0.71–1.15×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Opus 4.7 (max) / TasteValExperimental compute multiplier0.50×Hierarchical bootstrap: 95% CI 0.27–0.95×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Opus 4.7 (max) / TasteValNormalized final performance multiplier0.75×Hierarchical bootstrap: 95% CI 0.42–1.09×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-5.5 (xhigh) / TasteValExperimental compute multiplier0.47×Hierarchical bootstrap: 95% CI 0.24–1.11×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-5.5 (xhigh) / TasteValNormalized final performance multiplier0.84×Hierarchical bootstrap: 95% CI 0.55–1.13×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Kimi K3 (reasoning on) / TasteValExperimental compute multiplier0.45×Hierarchical bootstrap: 95% CI 0.31–0.71×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Kimi K3 (reasoning on) / TasteValNormalized final performance multiplier0.89×Hierarchical bootstrap: 95% CI 0.71–1.12×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Opus 4.6 (max) / TasteValExperimental compute multiplier0.40×Hierarchical bootstrap: 95% CI 0.24–0.85×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Opus 4.6 (max) / TasteValNormalized final performance multiplier0.78×Hierarchical bootstrap: 95% CI 0.49–1.14×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-5.2 (high) / TasteValExperimental compute multiplier0.28×Hierarchical bootstrap: 95% CI 0.16–0.54×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-5.2 (high) / TasteValNormalized final performance multiplier0.70×Hierarchical bootstrap: 95% CI 0.39–0.98×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-5.4 (xhigh) / TasteValExperimental compute multiplier0.21×Hierarchical bootstrap: 95% CI 0.12–0.37×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-5.4 (xhigh) / TasteValNormalized final performance multiplier0.60×Hierarchical bootstrap: 95% CI 0.36–0.84×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Opus 4.5 (24k thinking budget) / TasteValExperimental compute multiplier0.12×Hierarchical bootstrap: 95% CI 0.05–0.24×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Opus 4.5 (24k thinking budget) / TasteValNormalized final performance multiplier0.37×Hierarchical bootstrap: 95% CI 0.11–0.79×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-5 (high) / TasteValExperimental compute multiplier0.12×Hierarchical bootstrap: 95% CI 0.06–0.21×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-5 (high) / TasteValNormalized final performance multiplier0.53×Hierarchical bootstrap: 95% CI 0.22–0.73×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—Kimi K2.5 (max, 96k output) / TasteValExperimental compute multiplier0.09×Hierarchical bootstrap: 95% CI 0.04–0.19×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—Kimi K2.5 (max, 96k output) / TasteValNormalized final performance multiplier0.23×Hierarchical bootstrap: 95% CI 0.06–0.69×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—o3 (high) / TasteValExperimental compute multiplier0.08×Hierarchical bootstrap: 95% CI 0.04–0.16×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—o3 (high) / TasteValNormalized final performance multiplier0.43×Hierarchical bootstrap: 95% CI 0.18–0.61×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—o1 (high) / TasteValExperimental compute multiplier0.07×Hierarchical bootstrap: 95% CI 0.03–0.14×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—o1 (high) / TasteValNormalized final performance multiplier0.22×Hierarchical bootstrap: 95% CI 0.06–0.48×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-4o (gpt-4o-2024-05-13) / TasteValExperimental compute multiplier0.06×Hierarchical bootstrap: 95% CI 0.02–0.23×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-4o (gpt-4o-2024-05-13) / TasteValNormalized final performance multiplier0.29×Hierarchical bootstrap: 95% CI 0.09–0.59×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported
—GPT-4 (gpt-4-0613) / TasteValExperimental compute multiplier0.03×Hierarchical bootstrap: 95% CI 0.01–0.08×2026-10-05tasteval-v1-compute-multiplier-protocolbenchmark author reported
—GPT-4 (gpt-4-0613) / TasteValNormalized final performance multiplier0.22×Hierarchical bootstrap: 95% CI 0.06–0.55×2026-10-05tasteval-v1-performance-multiplier-protocolbenchmark author reported

Version lineage

Reference points

human baseline

1 · ×

Human-reference units derived from the recruited experts’ per-task best-of-run compute envelopes. Model scores use the paper’s chained-reference calculation; the envelope need not be one person’s single trajectory. Normalized to 1×.

What this reference means: measured baseline · Where it applies: established

Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.

Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.

The 1× normalization applies to each evaluated task subset; model aggregates do not all cover the same tasks.

Crossing this reference does not establish general research superiority or RSI.

human baseline

1 · ×

Best recruited human final score per task, using the same fixed Coder and compute/time caps; normalized to 1×.

What this reference means: measured baseline · Where it applies: established

Reference data come from 24 recruited experts, at least two per task, not the best researchers worldwide. Compute uses per-task best-of-run envelopes; final performance uses the best recruited human final score per task. Humans and models share the Coder and compute/time budgets; inference and salary costs are not matched.

Opus 5.5 and Fable 5.1 were scored on seven tasks; one refused task was excluded as missing data, not failure or zero. The other main-cohort models were scored on eight.

The 1× normalization applies to each evaluated task subset; model aggregates do not all cover the same tasks.

Crossing this reference does not establish general research superiority or RSI.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation servicepartial

Evidence in source charts

Official sources