benchmark

PostTrainBench

Tests how well AI agents improve small AI models, with ten hours and one H100 GPU per training run.

AI-R&D capabilityAI-system improvementEvaluation integrity

What this measure tells us

Tests bounded AI post-training research and experimental iteration.

  • Each ten-hour run optimizes one base model on one benchmark; the headline combines 28 target configurations.
  • Target-evaluation scores do not establish broad transfer or retained improvement to the research agent.
  • Different agents, seeds, compliance decisions and historical/rerun mixes can affect the aggregate.

Results by version

PostTrainBench v1.1 · September 2026 results

2026-07-28 · 28 tasks

Each result combines four base models and seven tests (28 target combinations per seed). These include BFCL v3 exec_simple, ArenaHard v2 creative writing, and a selected 245-question HealthBench-Easy set, rather than those entire benchmarks. The v1.1 release kept the original benchmark tasks and final evaluation. Results combine original runs that passed updated checks and reruns.

Metric definitions

Weighted downstream benchmark score (percent): Reported target-evaluation performance after a bounded post-training run, including official base-score fallback for flagged/missing cells. Unlike PostTrainBench Lite, this is not normalized reward.

Published results · PostTrainBench v1.1 · September 2026 results
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Fable 5 + Opus 4.8 Max GPQA fallbackWeighted downstream benchmark score41.79%Standard deviation across score seeds: ± 1.7% across 2 seeds2026-09-18ptb-v11-fable-5-protocolbenchmark author reported
—GLM 5.2, Max reasoning / Claude CodeWeighted downstream benchmark score31.70%Standard deviation across score seeds: ± 2.1% across 3 seeds2026-09-18ptb-v11-glm-5-2-protocolbenchmark author reported
—GPT-5.4, High reasoning / Codex CLIWeighted downstream benchmark score19.00%Standard deviation across score seeds: ± 3.0% across 3 seeds2026-09-18ptb-v11-gpt-5-4-high-protocolbenchmark author reported
—GPT-5.5, xHigh reasoning / Codex CLIWeighted downstream benchmark score27.23%Standard deviation across score seeds: ± 0.6% across 2 seeds2026-09-18ptb-v11-gpt-5-5-xhigh-protocolbenchmark author reported
—GPT-5.6 Sol, Max reasoning / Codex CLIWeighted downstream benchmark score36.23%Standard deviation across score seeds: ± 0.2% across 2 seeds2026-09-18ptb-v11-gpt-5-6-sol-protocolbenchmark author reported
—Gemini 3.1 Pro / OpenCodeWeighted downstream benchmark score21.99%Standard deviation across score seeds: ± 1.8% across 3 seeds2026-09-18ptb-v11-gemini-3-1-pro-protocolbenchmark author reported
—Grok 4.5, High reasoning / Cursor CLIWeighted downstream benchmark score23.45%Standard deviation across score seeds: ± 0.1% across 2 seeds2026-09-18ptb-v11-grok-4-5-high-protocolbenchmark author reported
—Kimi K3, 1M context / Claude CodeWeighted downstream benchmark score31.96%Standard deviation across score seeds: ± 0.3% across 3 seeds2026-09-18ptb-v11-kimi-k3-protocolbenchmark author reported
—Claude Opus 4.7, xHigh reasoning / Claude CodeWeighted downstream benchmark score28.56%Standard deviation across score seeds: ± 1.7% across 3 seeds2026-09-18ptb-v11-opus-4-7-protocolbenchmark author reported
—Claude Opus 4.8, High reasoning / Claude CodeWeighted downstream benchmark score33.84%Standard deviation across score seeds: ± 3.6% across 2 seeds2026-09-18ptb-v11-opus-4-8-protocolbenchmark author reported
—Claude Opus 4.8, Max reasoning / Claude CodeWeighted downstream benchmark score32.90%Standard deviation across score seeds: ± 5.8% across 2 seeds2026-09-18ptb-v11-opus-4-8-max-protocolbenchmark author reported
—Claude Opus 5 / Claude CodeWeighted downstream benchmark score35.04%Standard deviation across score seeds: ± 1.4% across 2 seeds2026-09-18ptb-v11-opus-5-protocolbenchmark author reported
—Locus (Intology, powered by Opus 5)Weighted downstream benchmark score45.58%Standard deviation across score seeds: ± 0.8% across 3 seeds2026-09-18ptb-v11-locus-protocolbenchmark author reported

Version lineage

Reference points

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Related benchmarks

Official sources