report · 2025-04-16

METR evaluation of o3 and o4-mini

METR

Why it matters here

An independent evaluation of research-task performance, with a separate five-task RE-Bench comparison.

AI-R&D capability

What the charts show

Research performance varies across tasks and evaluation setups

On five RE-Bench tasks, o3 with AIDE scored below o1-preview and Claude 3.7 Sonnet. o4-mini led, largely through one kernel-optimization task.

The chart compares four two-hour AI attempts with one eight-hour human attempt per task. It shows normalized scores with 95% bootstrap confidence intervals. Identified reward hacking counts as failure.

View original chart and methods ↗
Details about METR’s preliminary evaluation of OpenAI’s o3 and o4-mini · Executive Summary; RE-Bench methods; Performance on subset of RE-Bench (95% CI), rebench_bar.png

How to interpret this chart
  • Five-task subset; keep separate from the original seven-task study.
  • Simple scaffolds and limited testing may understate capability. o3/o4-mini used AIDE; Claude 3.7 used Modular.
  • Best-of-four AI attempts differ from one sustained human attempt. The separate 32-hour result is not this chart.
  • No exact values were transcribed from plotted positions.

What to keep in mind

  • This report uses a different task subset and setup from the original RE-Bench study.