Evidence record
Internal Research Debugging Evaluation
Astra card snapshot, September 2026 · GPT-6 Sol · 2026-09-03
Correction or update linkedIndexed four exact comparator labels from the Astra-card debugging figure; this is historical evidence, not a new model release.
Reported result · numeric
64.20%
Mean rubric reward · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardAppendix Figure 84, Internal Research Debugging, original labeled PNGOriginal source ↗
Evaluation setup
| aggregate method | Mean rubric reward (%) |
|---|---|
| comparability caveats | Source describes 41 bugs and six alignment-auditing tasks, but does not specify which contribute to the Figure 84 mean. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Internal task set and rubric are not public; no uncertainty intervals are printed.
Limitations
- Figure labels are mean rubric rewards; do not equate with older-card median wording.
- Internal suite; no cross-lab percentage comparison.
- Not a demonstration of recursive improvement.