Evidence record
PostTrainBench
PostTrainBench v1.2 · September 30 release · Claude Opus 5 / Claude Code · 2026-09-30
Reported result · numeric
35.66%
Standard deviation across score seeds: ± 2.62 percentage points across 2 score seeds
Weighted downstream benchmark score · percent
benchmark author reported · extraction review: agent checked
Sources
Published 2026-09-30
Evaluation setup
| split | AIME 2025, GPQA Main, GSM8K, HumanEval, ArenaHard creative writing and HealthBench-Easy; BFCL excluded. |
|---|---|
| task count | 24 |
| task snapshot | v1.2 score bundle at website commit d47f16ec4b7d68c7048e4dbbc154984980985bfa |
| run count | 2 |
| selection rule | Source-reported compliant/retrospectively rescored results. Flagged runs receive base scores; two-of-three contamination verdict with manual review. Exact per-run audit decisions not independently rechecked. |
| aggregate method | Mean over four base models of each weighted six-benchmark score; exact weights in the pinned v1.2 data, then reported mean across score seeds. |
| wall clock budget | Up to ten hours for each base-model × target-benchmark run |
| hardware | One H100 GPU per run |
| tool access | Native CLI, terminal, web, evaluator, shared n-gram decontamination checker |
| internet access | true |
| scaffold | Claude Code |
| evaluator version | PostTrainBench v1.2: isolated HumanEval tests; remote-code loading disabled for ArenaHard/HealthBench. Final evaluations use five, three or one evaluation seeds depending on target; no inference that aggregate n is this repetition count. |
| human intervention | No user interaction during each agent run under the benchmark prompt |
| task exclusions | BFCL removed. Six remaining target tests, including the writing and HealthBench subsets rather than entire broad benchmark families. |
| contamination concerns | Contamination requires two of three judge flags; API-use and benchmark-lookup judges run once. Judges use GPT-5.6 Terra; flagged runs reviewed manually. This does not prove all unflagged runs clean. |
| comparability caveats | Different evaluation rules and weights from v1.1; no continuous score trajectory across versions.; Aggregate score seeds differ from the repeated final-model evaluations; the published SD is not a confidence interval.; Retrospective rescoring includes historical runs; no claim all agents reran under identical new instructions. |
- PostTrainBench v1.2: More reliable scores and verdicts↗
- PostTrainBench v1.2 scores at website commit d47f16e↗
- PostTrainBench agent identities and scaffolds at website commit d47f16e↗
- PostTrainBench score-bundle generation at website commit d47f16e↗
Not reported: attempts per task, token budget, training budget, inference budget, monetary cost, filtering.
Comparability
Not compared with other results.
Limitations
- Different evaluation rules and weights from v1.1; no continuous score trajectory across versions.
- Aggregate score seeds differ from the repeated final-model evaluations; the published SD is not a confidence interval.
- Retrospective rescoring includes historical runs; no claim all agents reran under identical new instructions.
- Version 1.2 combines four base models and six target tests; historical v1.1 uses seven. Scores across these versions are not directly comparable.
- Google’s Gemini 4 Argon report uses a separate v1.1/OpenCode setup; keep those results distinct from the benchmark-author leaderboard.
- Each run optimizes one assigned model for one target test. A high aggregate does not establish broad transfer or improvement of the research agent itself.
- Fable 5.1 includes five GPQA run cells from Opus 5; the result is a mixed-system aggregate.