Evidence record

PostTrainBench Lite

Astra card snapshot, September 2026 · GPT-6 Astra · 2026-09-03

Correction or update linkedReplaced an incorrect missing-result label with the source’s published budget-curve evidence. Exact point values remain unavailable in the checked source.

Reported result · not reported

not reported

Gain toward instruction-tuned reference performance · normalized score

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-03

Evaluation setup

task count12
aggregate methodclamp((trial − base)/(instruct − base), 0, 1); cheating trials assigned zero
wall clock budget5 hours
hardwareOne H100 GPU
internet accesstrue

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • No exact chart value transcribed; missing is not zero.
  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.