Evidence record

MLE-bench

Revised, Astra card snapshot · GPT-6 Astra · 2026-09-03

Correction or update linkedTranscribed the explicitly labelled Revised MLE-bench score; preserved the separate historical medal-rate series.

Reported result · not reported

not reported

Mean percentile against reference solutions · percent

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-03

Evaluation setup

task count72
selection ruleUp to three leaderboard submissions
aggregate methodPercentile rank against test-time-compute reference distribution

Not reported: split, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • Numeric chart not transcribed; revised metric cannot be joined to medal rates.
  • Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
  • Public competition history creates contamination risk.