benchmark

MLE-bench

ML competition engineering under bounded compute.

AI-R&D capability

What this measure tells us

Measures practical training and data-science work.

  • Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
  • Public competition history creates contamination risk.

Results by version

Revised, Astra card snapshot

2026-09-03 · 72 tasks

Replaces tasks and scores percentile rank against a generated reference distribution; up to three leaderboard submissions.

Metric definitions

Mean percentile against reference solutions (percent): Percentile rank against a generated distribution of strong solutions; not the percentage of competitions earning a medal.

Published results · Revised, Astra card snapshot
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—GPT-6 AstraMean percentile against reference solutions93.80%2026-09-03mle-revised-maxlab reported
—GPT-5.6 SolMean percentile against reference solutions87.18%2026-09-03mle-revised-maxlab reported

Original paper v1 (2024)

2024-10-09 · 75 tasks

Mean any-medal fraction per attempt; not pass@k.

Metric definitions

Any medal rate (percent): A task-specific measurement; not a percentage of RSI achieved.

MLE-bench · Any medal rate

direct comparison · Point estimates; uncertainty is shown in the table.

Different seed counts affect precision; no significance inference from rounded means.

MLE-bench: Any medal rateCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.01020percent2024-10-09: 8.7 ± 0.5% · GPT-4o / AIDE · Standard error18.72024-10-09: 7.6 ± 1.8% · Claude 3.5 Sonnet / AIDE · Standard error27.62024-10-09: 3.0 ± 1.0% · Llama 3.1 405B / AIDE · Standard error332024-10-09: 16.9 ± 1.1% · o1-preview / AIDE · Standard error416.9
Published results · Original paper v1 (2024)
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
4o1-preview / AIDEAny medal rate16.9 ± 1.1%Standard error2024-10-09mle-o1-pbenchmark author reported
1GPT-4o / AIDEAny medal rate8.7 ± 0.5%Standard error2024-10-09mle-4o-pbenchmark author reported
3Llama 3.1 405B / AIDEAny medal rate3.0 ± 1.0%Standard error2024-10-09mle-llama-pbenchmark author reported
2Claude 3.5 Sonnet / AIDEAny medal rate7.6 ± 1.8%Standard error2024-10-09mle-claude-pbenchmark author reported
—GPT-4o / MLABAny medal rate0.8 ± 0.5%Standard error2024-10-09mle-4o-mlab-pbenchmark author reported
—GPT-4o / OpenHandsAny medal rate4.4 ± 1.4%Standard error2024-10-09mle-4o-openhands-pbenchmark author reported

Version lineage

Reference points

No reference points are indexed for this measure.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources