original paper

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering (v1)

OpenAI · Published 2024-10-09

Read the original document ↗

Evidence from this source

Results and research linked to this document
TypeEvidence
ResultMLE-bench · o1-preview / AIDE
Source details

Section 3.1, Table 2

ResultMLE-bench · GPT-4o / AIDE
Source details

Section 3.1, Table 2

ResultMLE-bench · Llama 3.1 405B / AIDE
Source details

Section 3.1, Table 2

ResultMLE-bench · Claude 3.5 Sonnet / AIDE
Source details

Section 3.1, Table 2

ResultMLE-bench · GPT-4o / MLAB
Source details

Section 3.1, Table 2, Scaffolding and Models experiments, Any Medal column

ResultMLE-bench · GPT-4o / OpenHands
Source details

Section 3.1, Table 2, Scaffolding and Models experiments, Any Medal column

BenchmarkMLE-bench
Source details

Sections 2–3; Table 2

Paper / reportMLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Source details

Primary publication source

Document tracking details
Source ID
src-mle
Source type
Primary
Original URL
https://arxiv.org/html/2410.07095v1
Last updated by publisher
Not reported
Access status
available
Retrieved
2026-09-26T04:47:27Z
Document fingerprint
ecfe27c79181a88c1358822316198b4e085441c2e26379b3c6ef49312918bf7d
Based on original bytes. Used to detect changes to the source.

Availability checks

  • 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.