preprint · 2026-03-09
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Why it matters here
PostTrainBench v1.1 tests AI agents improving assigned small base models under a ten-hour, one-H100 budget per target task. The September 2026 leaderboard snapshot reports 13 aggregate agent/configuration rows across four base models and seven weighted tests; the top external Locus result is 45.58% across three seeds. Official instruct models score 51.14% but use larger training budgets and are a contextual reference. Version 1.1 applies stronger integrity reviews and includes historical compliant runs and reruns.
AI-R&D capabilityAI-system improvementEvaluation integrity
Related benchmarks and indicators
What to keep in mind
- The paper’s 23.2% best-agent value refers to the March v1 release, not current v1.1.
- Fable 5’s displayed aggregate borrows Opus 4.8 Max GPQA cells after refusal; it is a reported mixed-system aggregate.
- The pinned collector substitutes untrained base scores for missing/broken runs and contamination/API flags; a benchmark-lookup flag blocks aggregation for investigation. Legacy results without comparable v1.1 traces may be omitted.
- The benchmark targets individual evaluations; high weighted scores need not imply general transferable model improvement.
- The pinned run prompt bars direct calls to hosted external model APIs; the API judge permits locally served model inference under stated conditions, so do not call all external-model distillation forbidden.
- The weighted suite includes selected BFCL, ArenaHard and HealthBench subsets, not each full benchmark; the independent Epoch review discusses this scope and judgment limits.
Related evidence
PostTrainBench · PostTrainBench v1.1 · September 2026 results41.79%PostTrainBench · PostTrainBench v1.1 · September 2026 results31.70%PostTrainBench · PostTrainBench v1.1 · September 2026 results19.00%PostTrainBench · PostTrainBench v1.1 · September 2026 results27.23%PostTrainBench · PostTrainBench v1.1 · September 2026 results36.23%PostTrainBench · PostTrainBench v1.1 · September 2026 results21.99%PostTrainBench · PostTrainBench v1.1 · September 2026 results23.45%PostTrainBench · PostTrainBench v1.1 · September 2026 results31.96%PostTrainBench · PostTrainBench v1.1 · September 2026 results28.56%PostTrainBench · PostTrainBench v1.1 · September 2026 results33.84%PostTrainBench · PostTrainBench v1.1 · September 2026 results32.90%PostTrainBench · PostTrainBench v1.1 · September 2026 results35.04%PostTrainBench · PostTrainBench v1.1 · September 2026 results45.58%