Astra card snapshot, September 2026
2026-09-03Disclosure snapshot. The source describes 41 research bugs and six alignment-auditing tasks, but does not explicitly identify the plotted mean’s scored denominator.
Metric definitions
Mean rubric reward (percent): A task-specific measurement; not a percentage of RSI achieved.
Internal Research Debugging Evaluation · Mean rubric reward
limited comparison
Internal task set and rubric are not public; no uncertainty intervals are printed.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| 5 | GPT-6 Astra | Mean rubric reward | 78.05% | 2026-09-03 | openai-research-debugging-protocol | lab reported | |
| 4 | GPT-6 Sol | Mean rubric reward | 64.20% | 2026-09-03 | openai-research-debugging-protocol | lab reported | |
| 2 | GPT-5.6 Sol | Mean rubric reward | 68.32% | 2026-09-03 | openai-research-debugging-protocol | lab reported | |
| 3 | GPT-6 Luna | Mean rubric reward | 46.62% | 2026-09-03 | openai-research-debugging-protocol | lab reported | |
| 1 | GPT-5.6 Luna | Mean rubric reward | 50.8% | 2026-09-03 | openai-research-debugging-protocol | lab reported |
No results match these filters for this version.