Evidence record

PostTrainBench Lite

Astra card snapshot, September 2026 · GPT-6 Astra · 2026-09-03

Correction or update linkedReplaced an incorrect missing-result label with the source’s published budget-curve evidence. Exact point values remain unavailable in the checked source.

Reported result · qualitative

Results appear in a chart at several output-token budgets; exact point values are not printed.

Gain toward instruction-tuned reference performance · normalized score

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-03

Evaluation setup

task count12
aggregate methodclamp((trial − base)/(instruct − base), 0, 1); cheating trials assigned zero
wall clock budget5 hours
hardwareOne H100 GPU
internet accesstrue

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • Checked official HTML, original PNG and report PDF; no labelled point table or downloadable structured results found. Values have not been estimated from chart heights.
  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.