Evidence record
Automated weak-to-strong research study
April 2026 chat-preference study · Human-tuned weak-to-strong method · 2026-04
Reported result · numeric
0.23 PGR
Performance gap recovered · PGR
lab reported · extraction review: agent checked
Sources
Published 2026-04
Evaluation setup
| split | Same chat preference held-out test, ID+OOD |
|---|---|
| selection rule | Strongest of four manually tuned literature methods plus a zero-shot baseline |
| wall clock budget | Two authors spent seven days tuning baselines |
| human intervention | Manual research and hyperparameter tuning by two authors |
| comparability caveats | Human time and AAR compute/API budgets are not matched; no causal speedup can be calculated. |
Not reported: task count, task snapshot, attempts per task, run count, aggregate method, token budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, task exclusions, contamination concerns.
Comparability
limited comparison
- Research time and cost are not matched.
- AAR repeatedly queried evaluation scores; its held-out test was not untouched.
Limitations
- This is the human-developed method result, not an AI-agent run.
- Two authors spent seven days; resource budgets are not matched.
- Not a general alignment or AI-research leaderboard.
- Agents repeatedly received test-evaluation scores during optimization; the test was not untouched.
- New research directions and evaluation infrastructure were supplied by humans; production-scale transfer gave no clear gain.