Evidence record
PostTrainBench
PostTrainBench v1.1 / Google OpenCode evaluation, October 2026 · Claude Opus 5.5 / Google OpenCode evaluation · 2026-10
Reported result · numeric
49.3%
Weighted downstream benchmark score · percent
lab reported · extraction review: agent checked
Sources
Published 2026-10
Evaluation setup
| aggregate method | Weighted aggregate across four base models and seven benchmarks; exact weights and per-cell results not restated. |
|---|---|
| wall clock budget | 10 hours |
| hardware | One NVIDIA H100 GPU |
| scaffold | OpenCode for all models |
| evaluator version | PostTrainBench v1.1; precise code revision not reported |
| comparability caveats | All four results were computed by Google using OpenCode; do not merge them with the benchmark-author leaderboard or other harnesses.; Model-specific reasoning effort, seeds, exact task snapshot, weights and contamination/audit outcomes are not fully specified. No uncertainty or statistical significance inferred.; General single-attempt prose is not treated as an exact count of post-training seeds or per-task attempts. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Google self-computed result; separate OpenCode evaluation from the benchmark-author leaderboard.
- Weighted downstream score, not percent RSI achieved or percent improvement over the starting model.
- No seeds, per-cell results or uncertainty reported here; no statistical significance inferred.
- Version 1.2 combines four base models and six target tests; historical v1.1 uses seven. Scores across these versions are not directly comparable.
- Google’s Gemini 4 Argon report uses a separate v1.1/OpenCode setup; keep those results distinct from the benchmark-author leaderboard.
- Each run optimizes one assigned model for one target test. A high aggregate does not establish broad transfer or improvement of the research agent itself.
- Fable 5.1 includes five GPQA run cells from Opus 5; the result is a mixed-system aggregate.