report · 2026-04

Automated Weak-to-Strong Researcher

Why it matters here

Nine Opus 4.6 research agents explored weak-to-strong supervision methods for small models over five days and 800 cumulative agent-hours. On the chat-preference testbed, their best result recovered 0.97 of the performance gap between weak- and strong-supervision baselines, versus 0.23 for a human-tuned method. Agents repeatedly received evaluation scores during optimization; this was not an untouched final test. Transfer to math/code varied and a production adaptation was within noise.

Read original source ↗

AI-R&D capabilityAI-system improvement

What the charts show

Distinct directions improved weak-to-strong research progress

Nine AI research agents assigned different research directions found stronger methods sooner than nine agents working without assigned directions.

The vertical axis shows performance gap recovered (PGR): 0 matches the weaker teacher, and 1 matches a stronger student trained with correct labels. Time adds together the hours used by all agents. Lines track the best methods found so far; the gray square marks the strongest human-developed comparison method.

Original chart comparing directed and undirected nine-agent search trajectories, with individual evaluations, frontiers and a human-tuned baseline marker.
Original chart from Anthropic, reproduced unchanged. Enlarge chart ↗

View original chart and methods ↗
Automated Weak-to-Strong Researcher · Article §3.1, Figure 5 and caption

How to interpret this chart
  • No unprinted curve point was estimated.
  • Repeated test-score feedback shaped the frontier, so it is not untouched final-test evidence.
  • The two teams differ in human-seeded research directions.

Discovered methods transferred unevenly across tasks

One discovered method worked well on chat and math but less well on code. A second performed worse than the weak teacher on code.

The two methods were developed on chat preference and then tested on math and code. Their settings could be tuned, but the methods themselves were not changed. Performance gap recovered (PGR) measures improvement beyond the weak teacher toward the strong-supervision reference. Each dashed line represents that task’s best human-developed comparison method.

Original bar chart with PGR for two methods across chat, math and coding tasks, including a −0.12 code result.
Original chart from Anthropic, reproduced unchanged. Enlarge chart ↗

View original chart and methods ↗
Automated Weak-to-Strong Researcher · Article §3.4, Figure 8 and caption

How to interpret this chart
  • Only two selected methods were transferred.
  • Repeated chat evaluation during discovery; new task evaluations are a narrower transfer check.
  • PGR can be negative and is not bounded to 0–1.
Published values from the chart

Exact numbers printed on Anthropic Figure 8. Dashed human-tuned baselines vary by task; their math/code values are not transcribed from unlabelled positions.

Discovered methods transferred unevenly across tasks
Discovered methodChat preference PGRMath verification PGRCoding verification PGR
CCS + Self-Distill0.970.940.47
EM Posterior0.780.75−0.12

Related benchmarks and indicators

What to keep in mind

  • A human chose the problem, directions and evaluation infrastructure; human and agent research budgets were not matched.
  • Repeated evaluator feedback and documented reward hacks constrain interpretation of the headline result.
  • The research agents did not improve their own retained researcher model across generations.

Study details

Bounded AI R&D method search. The agents did not retain an improved version of their own researcher/optimizer, and no sustained recursive capability gain is established.

Methods and study design
Study design and reported evidence
What changedtraining_algorithm, training_data, other
What stayed fixedResearcher-agent model identity; Weak teacher and strong student model identities; Externally supplied research problem and PGR evaluator
Improved system used as optimizer laterno
Generations attempted / acceptednot reported / not reported
Held-out transferSource Figure 8: CCS + Self-Distill PGR chat 0.97/math 0.94/code 0.47; EM Posterior 0.78/0.75/−0.12. Method edits barred on transfer but hyperparameter tuning allowed. Production Sonnet 4.0 adaptation gave +0.5 points within noise.
Resource accountingNine parallel Opus 4.6 agents; five elapsed days; 800 cumulative agent-hours; about USD 18,000 compute/API cost.
Human contributionsHumans set the problem, assigned distinct research directions, built the testbeds and baselines, provided training helpers and a remote scoring API.
Author claimsAutomated research can be practical on outcome-gradable problems.

Related evidence