PostTrainBench v1.1 · September 2026 results
2026-07-28 · 28 tasksEach result combines four base models and seven tests (28 target combinations per seed). These include BFCL v3 exec_simple, ArenaHard v2 creative writing, and a selected 245-question HealthBench-Easy set, rather than those entire benchmarks. The v1.1 release kept the original benchmark tasks and final evaluation. Results combine original runs that passed updated checks and reruns.
Metric definitions
Weighted downstream benchmark score (percent): Reported target-evaluation performance after a bounded post-training run, including official base-score fallback for flagged/missing cells. Unlike PostTrainBench Lite, this is not normalized reward.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| — | Fable 5 + Opus 4.8 Max GPQA fallback | Weighted downstream benchmark score | 41.79%Standard deviation across score seeds: ± 1.7% across 2 seeds | 2026-09-18 | ptb-v11-fable-5-protocol | benchmark author reported | |
| — | GLM 5.2, Max reasoning / Claude Code | Weighted downstream benchmark score | 31.70%Standard deviation across score seeds: ± 2.1% across 3 seeds | 2026-09-18 | ptb-v11-glm-5-2-protocol | benchmark author reported | |
| — | GPT-5.4, High reasoning / Codex CLI | Weighted downstream benchmark score | 19.00%Standard deviation across score seeds: ± 3.0% across 3 seeds | 2026-09-18 | ptb-v11-gpt-5-4-high-protocol | benchmark author reported | |
| — | GPT-5.5, xHigh reasoning / Codex CLI | Weighted downstream benchmark score | 27.23%Standard deviation across score seeds: ± 0.6% across 2 seeds | 2026-09-18 | ptb-v11-gpt-5-5-xhigh-protocol | benchmark author reported | |
| — | GPT-5.6 Sol, Max reasoning / Codex CLI | Weighted downstream benchmark score | 36.23%Standard deviation across score seeds: ± 0.2% across 2 seeds | 2026-09-18 | ptb-v11-gpt-5-6-sol-protocol | benchmark author reported | |
| — | Gemini 3.1 Pro / OpenCode | Weighted downstream benchmark score | 21.99%Standard deviation across score seeds: ± 1.8% across 3 seeds | 2026-09-18 | ptb-v11-gemini-3-1-pro-protocol | benchmark author reported | |
| — | Grok 4.5, High reasoning / Cursor CLI | Weighted downstream benchmark score | 23.45%Standard deviation across score seeds: ± 0.1% across 2 seeds | 2026-09-18 | ptb-v11-grok-4-5-high-protocol | benchmark author reported | |
| — | Kimi K3, 1M context / Claude Code | Weighted downstream benchmark score | 31.96%Standard deviation across score seeds: ± 0.3% across 3 seeds | 2026-09-18 | ptb-v11-kimi-k3-protocol | benchmark author reported | |
| — | Claude Opus 4.7, xHigh reasoning / Claude Code | Weighted downstream benchmark score | 28.56%Standard deviation across score seeds: ± 1.7% across 3 seeds | 2026-09-18 | ptb-v11-opus-4-7-protocol | benchmark author reported | |
| — | Claude Opus 4.8, High reasoning / Claude Code | Weighted downstream benchmark score | 33.84%Standard deviation across score seeds: ± 3.6% across 2 seeds | 2026-09-18 | ptb-v11-opus-4-8-protocol | benchmark author reported | |
| — | Claude Opus 4.8, Max reasoning / Claude Code | Weighted downstream benchmark score | 32.90%Standard deviation across score seeds: ± 5.8% across 2 seeds | 2026-09-18 | ptb-v11-opus-4-8-max-protocol | benchmark author reported | |
| — | Claude Opus 5 / Claude Code | Weighted downstream benchmark score | 35.04%Standard deviation across score seeds: ± 1.4% across 2 seeds | 2026-09-18 | ptb-v11-opus-5-protocol | benchmark author reported | |
| — | Locus (Intology, powered by Opus 5) | Weighted downstream benchmark score | 45.58%Standard deviation across score seeds: ± 0.8% across 3 seeds | 2026-09-18 | ptb-v11-locus-protocol | benchmark author reported |
No results match these filters for this version.