Evidence record
MLE-bench
Revised, Astra card snapshot · GPT-6 Astra · 2026-09-03
Correction or update linkedTranscribed the explicitly labelled Revised MLE-bench score; preserved the separate historical medal-rate series.
Reported result · not reported
not reported
Mean percentile against reference solutions · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardSection 10.1.3.5Original source ↗
Evaluation setup
| task count | 72 |
|---|---|
| selection rule | Up to three leaderboard submissions |
| aggregate method | Percentile rank against test-time-compute reference distribution |
Not reported: split, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Numeric chart not transcribed; revised metric cannot be joined to medal rates.
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.