Evidence record

MLE-bench

Original paper v1 (2024) · GPT-4o / MLAB · 2024-10-09

Correction or update linkedIndexed original MLE-bench GPT-4o scaffold comparisons from the 2024 paper.

Reported result · numeric

0.8 ± 0.5%

Standard error

Any medal rate · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2024-10-09

Evaluation setup

task count75
run count3
aggregate methodMean across repeated attempts; one SEM
wall clock budget24 hours per run
hardware36 vCPUs, 440 GB RAM, one Nvidia A10 GPU
scaffoldMLAB
comparability caveatsDifferent scaffold from AIDE; three seeds versus 36 for GPT-4o/AIDE.

Not reported: split, task snapshot, attempts per task, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Scaffold comparison; seed count differs from AIDE.
  • Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
  • Public competition history creates contamination risk.