Why it matters here
Nine Opus 4.6 research agents explored weak-to-strong supervision methods for small models over five days and 800 cumulative agent-hours. On the chat-preference testbed, their best result recovered 0.97 of the performance gap between weak- and strong-supervision baselines, versus 0.23 for a human-tuned method. Agents repeatedly received evaluation scores during optimization; this was not an untouched final test. Transfer to math/code varied and a production adaptation was within noise.
Read original source ↗
AI-R&D capabilityAI-system improvement
Related benchmarks and indicators
What to keep in mind
- A human chose the problem, directions and evaluation infrastructure; human and agent research budgets were not matched.
- Repeated evaluator feedback and documented reward hacks constrain interpretation of the headline result.
- The research agents did not improve their own retained researcher model across generations.
Study details
Bounded AI R&D method search. The agents did not retain an improved version of their own researcher/optimizer, and no sustained recursive capability gain is established.
Methods and study design
Study design and reported evidence| What changed | training_algorithm, training_data, other |
|---|
| What stayed fixed | Researcher-agent model identity; Weak teacher and strong student model identities; Externally supplied research problem and PGR evaluator |
|---|
| Improved system used as optimizer later | no |
|---|
| Generations attempted / accepted | not reported / not reported |
|---|
| Held-out transfer | Source Figure 8: CCS + Self-Distill PGR chat 0.97/math 0.94/code 0.47; EM Posterior 0.78/0.75/−0.12. Method edits barred on transfer but hyperparameter tuning allowed. Production Sonnet 4.0 adaptation gave +0.5 points within noise. |
|---|
| Resource accounting | Nine parallel Opus 4.6 agents; five elapsed days; 800 cumulative agent-hours; about USD 18,000 compute/API cost. |
|---|
| Human contributions | Humans set the problem, assigned distinct research directions, built the testbeds and baselines, provided training helpers and a remote scoring API. |
|---|
| Author claims | Automated research can be practical on outcome-gradable problems. |
|---|