Evidence record

PostTrainBench

PostTrainBench v1.1 · September 2026 results · GPT-5.6 Sol, Max reasoning / Codex CLI · 2026-09-18

Reported result · numeric

36.23%

Standard deviation across score seeds: ± 0.2% across 2 seeds

Weighted downstream benchmark score · percent

benchmark author reported · extraction review: agent checked

Sources

Published 2026-09-18

Evaluation setup
splitBFCL v3 exec_simple; ArenaHard v2 creative-writing; HealthBench-Easy selected 245-question subset; GPQA Main; AIME 2025, GSM8K and HumanEval as named in the evaluation suite.
task count28
task snapshotCurrent v1.1 aggregate from website commit 22705452f80106da9ff841b4db24df7128bb9057
run count2
selection rulePer-cell compliant final_model; missing or broken runs and contamination/API-flagged cells use untrained base-model fallback. A benchmark-lookup flag blocks aggregation pending investigation.
aggregate methodWeighted seven-benchmark score per base model, then mean over four models; displayed mean and SD across score seeds
wall clock budgetUp to ten hours for each base-model × target-benchmark run
hardwareOne H100 GPU per run
tool accessNative CLI, terminal, web, evaluator, shared n-gram decontamination checker
internet accesstrue
scaffoldCodex CLI, Max reasoning
evaluator versionPostTrainBench v1.1 published results; Inspect respects final_model generation_config.json
human interventionNo user interaction during each agent run under the benchmark prompt
task exclusionsFull BFCL, full ArenaHard and full HealthBench are not evaluated; only the stated subsets contribute to the seven-test weighted score.
contamination concernsSpecialized v1.1 contamination/API-use/benchmark-lookup judges and model-identity checks; hosted external model APIs are prohibited but local inference servers are allowed under the run rules; historical unflagged v1 traces may remain
comparability caveatsThree earlier runs flagged for consulting PostTrainBench materials; current published aggregate follows team adjudication and base-score fallback.; Leaderboard v1.1 re-audits historical runs; do not interpret all scores as freshly run under one identical v1.1 agent instruction set.; The pinned run prompt prohibits direct calls to hosted external models for post-training; its API judge permits local inference servers under stated conditions.

Not reported: attempts per task, token budget, training budget, inference budget, monetary cost, filtering.

Comparability

Not compared with other results.

Limitations

  • Three earlier runs flagged for consulting PostTrainBench materials; current published aggregate follows team adjudication and base-score fallback.
  • The mean is across four base models and seven weighted target tests, not one agent run creating a generally improved model.
  • Seed SD is descriptive dispersion, not a confidence interval.
  • Each ten-hour run optimizes one base model on one benchmark; the headline combines 28 target configurations.
  • Target-evaluation scores do not establish broad transfer or retained improvement to the research agent.
  • Different agents, seeds, compliance decisions and historical/rerun mixes can affect the aggregate.