report · 2026-07-21

Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT

Why it matters here

METR proposes measuring optimization ability by the expenditure at which an agent’s improvement matches an estimated human improvement at the same budget. Its preliminary NanoGPT experiments separate raw scores, revalidated improvements and a maintainer’s estimate of useful contributions. The results illustrate why apparent progress, research cost and practical usefulness need separate checks.

Read original source ↗

AI-R&D capabilityAI-system improvementEvaluation integrity

What the charts show

Estimated optimization value depends on validation and the human cost reference

METR reports positive revalidated expenditure horizons for four of six models. Raw trajectories overstate progress; the potentially mergeable share is smaller than the full validated gain.

Each curve is one optimization trajectory from NanoGPT record #78. The horizon is the budget where agent and estimated human improvements match, using about $2,500 of human labor per 1% speedup. Agent cost includes inference and experiments; agents had 32 H100 GPUs across four nodes. Stars mark maintainer estimates, not confirmed merges.

Six panels show training time in seconds against expenditure, with raw and revalidated agent curves, an estimated human line, intersections, and two mergeability stars.
Unmodified original PNG from https://metr.org/assets/images/expenditure-horizon/expenditure-horizon.png; original CC-BY attribution retained. Enlarge chart ↗

View original chart and methods ↗
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT · Estimating agent expenditure horizons; expenditure-horizon.png; human-cost sensitivity table

How to interpret this chart
  • The human calibration relies on two contributor interviews and speculative effort estimates; it is not a matched human trial.
  • The $3,333 and $2,347 figure labels are estimates; the sensitivity table rounds them to $3,300 and $2,300.
  • Under $1,000 versus $10,000 per 1% human-cost assumptions, the article reports Opus-4.8 horizons of $120 versus $14,400 and GPT-5.5 horizons of $160 versus $9,400.
  • Noise, hardware differences, a partially AI-optimized starting checkpoint and possible contamination limit generalization. Record #12 contaminated runs are excluded from the main analysis.
  • Dollar-valued optimization ability is not a direct measure of recursive acceleration or actual economic savings.
Published values from the chart

Exact printed figure labels for the $2,500-per-1%-speedup human reference. Zero is explicitly printed, not substituted for missing data. These are estimated curve intersections, not API bills, human task durations, or a cross-protocol ranking.

Estimated optimization value depends on validation and the human cost reference
Model labelExpenditure horizon estimate (USD)
GPT 50
GPT 5.2840
GPT 5.52,347
Opus 4.10
Opus 4.5613
Opus 4.83,333

What to keep in mind

  • One narrow training-efficiency task, with one displayed trajectory per model and an uncertain human labor reference.
  • This is a July report indexed later, not an October model release or a measurement of frontier-wide AI R&D automation.
  • The tracker inspected original evidence; it did not reproduce METR’s validation or verify that proposed changes were merged.