reported capability index

Automated weak-to-strong research study

Tests whether AI research agents can find better ways to train a stronger model using guidance from a weaker model.

AI-R&D capabilityAI-system improvement

What this measure tells us

Shows bounded agent research and method improvement on a specified, externally scored AI-training problem.

  • Not a general alignment or AI-research leaderboard.
  • Agents repeatedly received test-evaluation scores during optimization; the test was not untouched.
  • New research directions and evaluation infrastructure were supplied by humans; production-scale transfer gave no clear gain.

Results by version

April 2026 chat-preference study

2026-04

Qwen1.5-0.5B-Chat weak teacher and Qwen3-4B-Base strong student; chat preference hill-climbing. PGR can be negative in cross-task transfer, so no universal [0,1] range is imposed.

Metric definitions

Performance gap recovered (PGR): 0 denotes no gain over the weak teacher; 1 matches the ground-truth-supervised strong student. Values can be negative. This is not research-automation percentage.

Automated weak-to-strong research study · Performance gap recovered

limited comparison

Research time and cost are not matched. AAR repeatedly queried evaluation scores; its held-out test was not untouched.

Automated weak-to-strong research study: Performance gap recoveredCategorical point plot with no connecting line. Author-reported PGR on the same chat-preference testbed and weak/strong student models.0.20.61PGRHuman reference · 0.232026-04: 0.97 PGR · Nine Claude Opus 4.6 research agents10.972026-04: 0.23 PGR · Human-tuned weak-to-strong method20.23
Published results · April 2026 chat-preference study
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
1Nine Claude Opus 4.6 research agentsPerformance gap recovered0.97 PGRAgents received evaluation scores throughout the search; this was not an untouched final test.2026-04automated-w2s-aar-protocollab reported
2Human-tuned weak-to-strong methodPerformance gap recovered0.23 PGR2026-04automated-w2s-human-protocollab reported

Version lineage

Reference points

human baseline

0.23 · PGR

Best manually tuned method among four literature baselines plus a zero-shot baseline; two authors worked seven days.

What this reference means: measured baseline · Where it applies: established

Researcher effort and agent compute/API budgets differ.

AAR repeatedly received test scores during method search.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public taskspartial
public codeyes
public evaluation serviceunknown

Evidence in source charts

Official sources