Evidence record

Automated weak-to-strong research study

April 2026 chat-preference study · Nine Claude Opus 4.6 research agents · 2026-04

Reported result · numeric

0.97 PGR

Agents received evaluation scores throughout the search; this was not an untouched final test.

Performance gap recovered · PGR

lab reported · extraction review: agent checked

Sources

Published 2026-04

Evaluation setup
splitChat preference ID+OOD held-out splits; repeated remote test-score feedback during optimization
run count9
selection ruleBest method selected from nine directed AAR research trajectories using repeated PGR submissions
wall clock budgetFive elapsed days; 800 cumulative agent-hours across nine parallel AARs
training budgetQwen1.5-0.5B-Chat weak teacher and Qwen3-4B-Base strong student; each fine-tune under two hours
monetary cost[object Object]
tool accessIndependent sandboxes; training helpers; shared forum/code; unlimited remote PGR submissions
scaffoldClaude Opus 4.6 AAR agents with human-specified ambiguous distinct research directions
human interventionHumans chose problem, directions, testbeds, baselines and evaluator; agents proposed, ran and shared experiments
contamination concernsRepeated test feedback makes the nominal held-out test effectively a validation set; source reports reward hacking
comparability caveatsA best-of-search result after repeated evaluation, not an untouched confirmatory test.; Agent cumulative hours, elapsed time and monetary cost are different resource measures.

Not reported: task count, task snapshot, attempts per task, aggregate method, token budget, hardware, inference budget, internet access, filtering, evaluator version, task exclusions.

Comparability

limited comparison

  • Research time and cost are not matched.
  • AAR repeatedly queried evaluation scores; its held-out test was not untouched.

Limitations

  • Agents repeatedly received evaluation scores during optimization; this was not an untouched final test.
  • The result is improvement in a small-model weak-to-strong testbed, not 97% automation of research.
  • Not a general alignment or AI-research leaderboard.
  • Agents repeatedly received test-evaluation scores during optimization; the test was not untouched.
  • New research directions and evaluation infrastructure were supplied by humans; production-scale transfer gave no clear gain.