Evidence record

Automated weak-to-strong research study

April 2026 chat-preference study · Human-tuned weak-to-strong method · 2026-04

Reported result · numeric

0.23 PGR

Performance gap recovered · PGR

lab reported · extraction review: agent checked

Sources

Published 2026-04

Evaluation setup
splitSame chat preference held-out test, ID+OOD
selection ruleStrongest of four manually tuned literature methods plus a zero-shot baseline
wall clock budgetTwo authors spent seven days tuning baselines
human interventionManual research and hyperparameter tuning by two authors
comparability caveatsHuman time and AAR compute/API budgets are not matched; no causal speedup can be calculated.

Not reported: task count, task snapshot, attempts per task, run count, aggregate method, token budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, task exclusions, contamination concerns.

Comparability

limited comparison

  • Research time and cost are not matched.
  • AAR repeatedly queried evaluation scores; its held-out test was not untouched.

Limitations

  • This is the human-developed method result, not an AI-agent run.
  • Two authors spent seven days; resource budgets are not matched.
  • Not a general alignment or AI-research leaderboard.
  • Agents repeatedly received test-evaluation scores during optimization; the test was not untouched.
  • New research directions and evaluation infrastructure were supplied by humans; production-scale transfer gave no clear gain.