Evidence record
MLE-bench
Original paper v1 (2024) · Claude 3.5 Sonnet / AIDE · 2024-10-09
Correction or update linkedCleared a metadata field that repeated the standard-error amplitude as if it were a confidence level; printed values and reported standard errors are unchanged.
Reported result · numeric
7.6 ± 1.8%
Standard error
Any medal rate · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2024-10-09
Evaluation setup
| task count | 75 |
|---|---|
| run count | 3 |
| aggregate method | Mean across repeated attempts; one SEM |
| wall clock budget | 24 hours per run |
| hardware | 36 vCPUs, 440 GB RAM, one Nvidia A10 GPU |
| scaffold | AIDE |
| comparability caveats | Number of seeds differs; comparison explicitly reported by authors. |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
direct comparison
- Different seed counts affect precision; no significance inference from rounded means.
Limitations
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.