Evidence record

PostTrainBench

PostTrainBench v1.1 / Google OpenCode evaluation, October 2026 · Claude Opus 5.5 / Google OpenCode evaluation · 2026-10

Reported result · numeric

49.3%

Weighted downstream benchmark score · percent

lab reported · extraction review: agent checked

Sources

Published 2026-10

Evaluation setup
aggregate methodWeighted aggregate across four base models and seven benchmarks; exact weights and per-cell results not restated.
wall clock budget10 hours
hardwareOne NVIDIA H100 GPU
scaffoldOpenCode for all models
evaluator versionPostTrainBench v1.1; precise code revision not reported
comparability caveatsAll four results were computed by Google using OpenCode; do not merge them with the benchmark-author leaderboard or other harnesses.; Model-specific reasoning effort, seeds, exact task snapshot, weights and contamination/audit outcomes are not fully specified. No uncertainty or statistical significance inferred.; General single-attempt prose is not treated as an exact count of post-training seeds or per-task attempts.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Google self-computed result; separate OpenCode evaluation from the benchmark-author leaderboard.
  • Weighted downstream score, not percent RSI achieved or percent improvement over the starting model.
  • No seeds, per-cell results or uncertainty reported here; no statistical significance inferred.
  • Version 1.2 combines four base models and six target tests; historical v1.1 uses seven. Scores across these versions are not directly comparable.
  • Google’s Gemini 4 Argon report uses a separate v1.1/OpenCode setup; keep those results distinct from the benchmark-author leaderboard.
  • Each run optimizes one assigned model for one target test. A high aggregate does not establish broad transfer or improvement of the research agent itself.
  • Fable 5.1 includes five GPQA run cells from Opus 5; the result is a mixed-system aggregate.