Evidence record
PostTrainBench
PostTrainBench v1.1 · September 2026 results · Claude Opus 4.7, xHigh reasoning / Claude Code · 2026-09-18
Reported result · numeric
28.56%
Standard deviation across score seeds: ± 1.7% across 3 seeds
Weighted downstream benchmark score · percent
benchmark author reported · extraction review: agent checked
Sources
Published 2026-09-18
Evaluation setup
| split | BFCL v3 exec_simple; ArenaHard v2 creative-writing; HealthBench-Easy selected 245-question subset; GPQA Main; AIME 2025, GSM8K and HumanEval as named in the evaluation suite. |
|---|---|
| task count | 28 |
| task snapshot | Current v1.1 aggregate from website commit 22705452f80106da9ff841b4db24df7128bb9057 |
| run count | 3 |
| selection rule | Per-cell compliant final_model; missing or broken runs and contamination/API-flagged cells use untrained base-model fallback. A benchmark-lookup flag blocks aggregation pending investigation. |
| aggregate method | Weighted seven-benchmark score per base model, then mean over four models; displayed mean and SD across score seeds |
| wall clock budget | Up to ten hours for each base-model × target-benchmark run |
| hardware | One H100 GPU per run |
| tool access | Native CLI, terminal, web, evaluator, shared n-gram decontamination checker |
| internet access | true |
| scaffold | Claude Code, xHigh reasoning |
| evaluator version | PostTrainBench v1.1 published results; Inspect respects final_model generation_config.json |
| human intervention | No user interaction during each agent run under the benchmark prompt |
| task exclusions | Full BFCL, full ArenaHard and full HealthBench are not evaluated; only the stated subsets contribute to the seven-test weighted score. |
| contamination concerns | Specialized v1.1 contamination/API-use/benchmark-lookup judges and model-identity checks; hosted external model APIs are prohibited but local inference servers are allowed under the run rules; historical unflagged v1 traces may remain |
| comparability caveats | Current v1.1 leaderboard aggregate.; Leaderboard v1.1 re-audits historical runs; do not interpret all scores as freshly run under one identical v1.1 agent instruction set.; The pinned run prompt prohibits direct calls to hosted external models for post-training; its API judge permits local inference servers under stated conditions. |
- PostTrainBench website scores.json at commit 2270545↗
- PostTrainBench v1.1 leaderboard and methodology↗
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?↗
- Hardening PostTrainBench against reward hacking↗
- PostTrainBench agent run prompt and rules↗
- PostTrainBench API use judge prompt↗
- PostTrainBench judge policy and score effects↗
- PostTrainBench score collection and fallback code↗
Not reported: attempts per task, token budget, training budget, inference budget, monetary cost, filtering.
Comparability
Not compared with other results.
Limitations
- Current v1.1 leaderboard aggregate.
- The mean is across four base models and seven weighted target tests, not one agent run creating a generally improved model.
- Seed SD is descriptive dispersion, not a confidence interval.
- Each ten-hour run optimizes one base model on one benchmark; the headline combines 28 target configurations.
- Target-evaluation scores do not establish broad transfer or retained improvement to the research agent.
- Different agents, seeds, compliance decisions and historical/rerun mixes can affect the aggregate.