Evidence record
Internal Research Debugging Evaluation
Astra card snapshot, September 2026 · GPT-6 Astra · 2026-09-03
Reported result · numeric
78.05%
Mean rubric reward · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardSection 10.1.3.1; Appendix Figure 84, Internal Research Debugging, original labeled PNGOriginal source ↗
Evaluation setup
| aggregate method | Mean rubric reward (%) |
|---|---|
| comparability caveats | Source describes 41 bugs and six alignment-auditing tasks, but does not specify which contribute to the Figure 84 mean. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Internal task set and rubric are not public; no uncertainty intervals are printed.
Limitations
- 41 debugging bugs; six additional alignment-auditing tasks are described separately.
- Internal suite; no cross-lab percentage comparison.
- Not a demonstration of recursive improvement.