Research performance varies across tasks and evaluation setups
On five RE-Bench tasks, o3 with AIDE scored below o1-preview and Claude 3.7 Sonnet. o4-mini led, largely through one kernel-optimization task.
The chart compares four two-hour AI attempts with one eight-hour human attempt per task. It shows normalized scores with 95% bootstrap confidence intervals. Identified reward hacking counts as failure.
View original chart and methods ↗
Details about METR’s preliminary evaluation of OpenAI’s o3 and o4-mini · Executive Summary; RE-Bench methods; Performance on subset of RE-Bench (95% CI), rebench_bar.png
How to interpret this chart
- Five-task subset; keep separate from the original seven-task study.
- Simple scaffolds and limited testing may understate capability. o3/o4-mini used AIDE; Claude 3.7 used Modular.
- Best-of-four AI attempts differ from one sustained human attempt. The separate 32-hour result is not this chart.
- No exact values were transcribed from plotted positions.