preprint · 2026-03-09

PostTrainBench: Can LLM Agents Automate LLM Post-Training?

Why it matters here

PostTrainBench v1.1 tests AI agents improving assigned small base models under a ten-hour, one-H100 budget per target task. The September 2026 leaderboard snapshot reports 13 aggregate agent/configuration rows across four base models and seven weighted tests; the top external Locus result is 45.58% across three seeds. Official instruct models score 51.14% but use larger training budgets and are a contextual reference. Version 1.1 applies stronger integrity reviews and includes historical compliant runs and reruns.

Read original source ↗

AI-R&D capabilityAI-system improvementEvaluation integrity

Related benchmarks and indicators

What to keep in mind

  • The paper’s 23.2% best-agent value refers to the March v1 release, not current v1.1.
  • Fable 5’s displayed aggregate borrows Opus 4.8 Max GPQA cells after refusal; it is a reported mixed-system aggregate.
  • The pinned collector substitutes untrained base scores for missing/broken runs and contamination/API flags; a benchmark-lookup flag blocks aggregation for investigation. Legacy results without comparable v1.1 traces may be omitted.
  • The benchmark targets individual evaluations; high weighted scores need not imply general transferable model improvement.
  • The pinned run prompt bars direct calls to hosted external model APIs; the API judge permits locally served model inference under stated conditions, so do not call all external-model distillation forbidden.
  • The weighted suite includes selected BFCL, ArenaHard and HealthBench subsets, not each full benchmark; the independent Epoch review discusses this scope and judgment limits.

Related evidence