Evidence record
Automated weak-to-strong research study
April 2026 chat-preference study · Nine Claude Opus 4.6 research agents · 2026-04
Reported result · numeric
0.97 PGR
Agents received evaluation scores throughout the search; this was not an untouched final test.
Performance gap recovered · PGR
lab reported · extraction review: agent checked
Sources
Published 2026-04
Evaluation setup
| split | Chat preference ID+OOD held-out splits; repeated remote test-score feedback during optimization |
|---|---|
| run count | 9 |
| selection rule | Best method selected from nine directed AAR research trajectories using repeated PGR submissions |
| wall clock budget | Five elapsed days; 800 cumulative agent-hours across nine parallel AARs |
| training budget | Qwen1.5-0.5B-Chat weak teacher and Qwen3-4B-Base strong student; each fine-tune under two hours |
| monetary cost | [object Object] |
| tool access | Independent sandboxes; training helpers; shared forum/code; unlimited remote PGR submissions |
| scaffold | Claude Opus 4.6 AAR agents with human-specified ambiguous distinct research directions |
| human intervention | Humans chose problem, directions, testbeds, baselines and evaluator; agents proposed, ran and shared experiments |
| contamination concerns | Repeated test feedback makes the nominal held-out test effectively a validation set; source reports reward hacking |
| comparability caveats | A best-of-search result after repeated evaluation, not an untouched confirmatory test.; Agent cumulative hours, elapsed time and monetary cost are different resource measures. |
Not reported: task count, task snapshot, attempts per task, aggregate method, token budget, hardware, inference budget, internet access, filtering, evaluator version, task exclusions.
Comparability
limited comparison
- Research time and cost are not matched.
- AAR repeatedly queried evaluation scores; its held-out test was not untouched.
Limitations
- Agents repeatedly received evaluation scores during optimization; this was not an untouched final test.
- The result is improvement in a small-model weak-to-strong testbed, not 97% automation of research.
- Not a general alignment or AI-research leaderboard.
- Agents repeatedly received test-evaluation scores during optimization; the test was not untouched.
- New research directions and evaluation infrastructure were supplied by humans; production-scale transfer gave no clear gain.