benchmark
MLE-bench ML competition engineering under bounded compute.
Results by version Benchmark version Revised, Astra card snapshot Original paper v1 (2024)
Metric All metrics Mean percentile against reference solutions · Revised, Astra card snapshot Any medal rate · Original paper v1 (2024)
Results to compare All comparison groups direct · mle-aide
Protocol All protocols mle-o1-p mle-4o-p mle-llama-p mle-claude-p mle-revised-p mle-revised-max mle-4o-mlab-p mle-4o-openhands-p
Revised, Astra card snapshot 2026-09-03 · 72 tasks Replaces tasks and scores percentile rank against a generated reference distribution; up to three leaderboard submissions.
Metric definitions Mean percentile against reference solutions (percent): Percentile rank against a generated distribution of strong solutions; not the percentage of competitions earning a medal.
Published results · Revised, Astra card snapshot Key System / organization Metric Reported result Date Protocol Verification Evidence — GPT-6 Astra Mean percentile against reference solutions 93.80% 2026-09-03 mle-revised-max lab reported — GPT-5.6 Sol Mean percentile against reference solutions 87.18% 2026-09-03 mle-revised-max lab reported
No results match these filters for this version.
Original paper v1 (2024) 2024-10-09 · 75 tasks Mean any-medal fraction per attempt; not pass@k.
Metric definitions Any medal rate (percent): A task-specific measurement; not a percentage of RSI achieved.
MLE-bench · Any medal rate direct comparison · Point estimates; uncertainty is shown in the table.
Different seed counts affect precision; no significance inference from rounded means.
MLE-bench: Any medal rate Categorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design. 0 10 20 percent 2024-10-09: 8.7 ± 0.5% · GPT-4o / AIDE · Standard error 1 8.7 2024-10-09: 7.6 ± 1.8% · Claude 3.5 Sonnet / AIDE · Standard error 2 7.6 2024-10-09: 3.0 ± 1.0% · Llama 3.1 405B / AIDE · Standard error 3 3 2024-10-09: 16.9 ± 1.1% · o1-preview / AIDE · Standard error 4 16.9 Plot hidden because a single protocol is selected; use the filtered result table below.
Published results · Original paper v1 (2024) Key System / organization Metric Reported result Date Protocol Verification Evidence 4 o1-preview / AIDE Any medal rate 16.9 ± 1.1% Standard error 2024-10-09 mle-o1-p benchmark author reported 1 GPT-4o / AIDE Any medal rate 8.7 ± 0.5% Standard error 2024-10-09 mle-4o-p benchmark author reported 3 Llama 3.1 405B / AIDE Any medal rate 3.0 ± 1.0% Standard error 2024-10-09 mle-llama-p benchmark author reported 2 Claude 3.5 Sonnet / AIDE Any medal rate 7.6 ± 1.8% Standard error 2024-10-09 mle-claude-p benchmark author reported — GPT-4o / MLAB Any medal rate 0.8 ± 0.5% Standard error 2024-10-09 mle-4o-mlab-p benchmark author reported — GPT-4o / OpenHands Any medal rate 4.4 ± 1.4% Standard error 2024-10-09 mle-4o-openhands-p benchmark author reported
No results match these filters for this version.