Evidence record

PostTrainBench

PostTrainBench v1.2 · September 30 release · GPT-6 Astra / Codex CLI, Max · 2026-09-30

Reported result · numeric

41.88%

Standard deviation across score seeds: ± 0.76 percentage points across 2 score seeds

Weighted downstream benchmark score · percent

benchmark author reported · extraction review: agent checked

Sources

Published 2026-09-30

Evaluation setup
splitAIME 2025, GPQA Main, GSM8K, HumanEval, ArenaHard creative writing and HealthBench-Easy; BFCL excluded.
task count24
task snapshotv1.2 score bundle at website commit d47f16ec4b7d68c7048e4dbbc154984980985bfa
run count2
selection ruleSource-reported compliant/retrospectively rescored results. Flagged runs receive base scores; two-of-three contamination verdict with manual review. Exact per-run audit decisions not independently rechecked.
aggregate methodMean over four base models of each weighted six-benchmark score; exact weights in the pinned v1.2 data, then reported mean across score seeds.
wall clock budgetUp to ten hours for each base-model × target-benchmark run
hardwareOne H100 GPU per run
tool accessNative CLI, terminal, web, evaluator, shared n-gram decontamination checker
internet accesstrue
scaffoldCodex CLI, Max reasoning
evaluator versionPostTrainBench v1.2: isolated HumanEval tests; remote-code loading disabled for ArenaHard/HealthBench. Final evaluations use five, three or one evaluation seeds depending on target; no inference that aggregate n is this repetition count.
human interventionNo user interaction during each agent run under the benchmark prompt
task exclusionsBFCL removed. Six remaining target tests, including the writing and HealthBench subsets rather than entire broad benchmark families.
contamination concernsContamination requires two of three judge flags; API-use and benchmark-lookup judges run once. Judges use GPT-5.6 Terra; flagged runs reviewed manually. This does not prove all unflagged runs clean.
comparability caveatsDifferent evaluation rules and weights from v1.1; no continuous score trajectory across versions.; Aggregate score seeds differ from the repeated final-model evaluations; the published SD is not a confidence interval.; Retrospective rescoring includes historical runs; no claim all agents reran under identical new instructions.

Not reported: attempts per task, token budget, training budget, inference budget, monetary cost, filtering.

Comparability

Not compared with other results.

Limitations

  • Different evaluation rules and weights from v1.1; no continuous score trajectory across versions.
  • Aggregate score seeds differ from the repeated final-model evaluations; the published SD is not a confidence interval.
  • Retrospective rescoring includes historical runs; no claim all agents reran under identical new instructions.
  • Version 1.2 combines four base models and six target tests; historical v1.1 uses seven. Scores across these versions are not directly comparable.
  • Google’s Gemini 4 Argon report uses a separate v1.1/OpenCode setup; keep those results distinct from the benchmark-author leaderboard.
  • Each run optimizes one assigned model for one target test. A high aggregate does not establish broad transfer or improvement of the research agent itself.
  • Fable 5.1 includes five GPQA run cells from Opus 5; the result is a mixed-system aggregate.