Evidence record
PostTrainBench Lite
Astra card snapshot, September 2026 · GPT-6 Astra · 2026-09-03
Correction or update linkedReplaced an incorrect missing-result label with the source’s published budget-curve evidence. Exact point values remain unavailable in the checked source.
Reported result · not reported
not reported
Gain toward instruction-tuned reference performance · normalized score
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardSection 10.1.3.4Original source ↗
Evaluation setup
| task count | 12 |
|---|---|
| aggregate method | clamp((trial − base)/(instruct − base), 0, 1); cheating trials assigned zero |
| wall clock budget | 5 hours |
| hardware | One H100 GPU |
| internet access | true |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- No exact chart value transcribed; missing is not zero.
- Internal suite; no cross-lab percentage comparison.
- Not a demonstration of recursive improvement.