Evidence record

Internal Research Debugging Evaluation

Astra card snapshot, September 2026 · GPT-6 Astra · 2026-09-03

Reported result · numeric

78.05%

Mean rubric reward · percent

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-03

Evaluation setup

aggregate methodMean rubric reward (%)
comparability caveatsSource describes 41 bugs and six alignment-auditing tasks, but does not specify which contribute to the Figure 84 mean.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Internal task set and rubric are not public; no uncertainty intervals are printed.

Limitations

  • 41 debugging bugs; six additional alignment-auditing tasks are described separately.
  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.