Evidence record

MLE-bench

Original paper v1 (2024) · o1-preview / AIDE · 2024-10-09

Correction or update linkedCleared a metadata field that repeated the standard-error amplitude as if it were a confidence level; printed values and reported standard errors are unchanged.

Reported result · numeric

16.9 ± 1.1%

Standard error

Any medal rate · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2024-10-09

Evaluation setup

task count75
run count16
aggregate methodMean across repeated attempts; one SEM
wall clock budget24 hours per run
hardware36 vCPUs, 440 GB RAM, one Nvidia A10 GPU
scaffoldAIDE
comparability caveatsNumber of seeds differs; comparison explicitly reported by authors.

Not reported: split, task snapshot, attempts per task, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

direct comparison

  • Different seed counts affect precision; no significance inference from rounded means.

Limitations

  • Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
  • Public competition history creates contamination risk.