report · 2026

When AI builds itself

Why it matters here

Choosing what to try next is part of doing research, not just writing code. Anthropic tested whether Claude could suggest a better direction at selected points where a researcher had gone off track. Newer models performed better on this test; it does not show that AI generally makes better research decisions than people.

Read original source ↗

AI-R&D capability

What the charts show

Choosing a better next step in a research problem

On selected research problems where a human had taken a wrong turn, Claude Mythos Preview suggested a better next step in 64% of cases, compared with 51% for Claude Opus 4.5. This suggests progress in deciding what to try next, beyond carrying out instructions.

Anthropic replayed 129 moments from January–March 2026 research sessions. A separate Claude, shown how each session ended, compared the proposed next step with the human’s original choice. These examples were deliberately selected because the human choice could be improved.

Original Anthropic chart comparing nine models on 129 selected research decisions. Better-choice shares range from 22% for Claude Haiku 3 to 64% for Claude Mythos Preview. Ties are shown separately. A 90% hindsight reference is marked as the practical ceiling.
Original chart by Anthropic. Labels and image retained without redrawing. Enlarge chart ↗

View original chart and methods ↗
When AI builds itself · Evidence from within Anthropic: research-next-step experiment and footnote 8

How to interpret this chart
  • The 129 examples were selected because the human next step had room for improvement; this is not a representative comparison with researchers.
  • A separate Claude judged proposals using the eventual session outcome. Proposals were not tested by running complete alternative research projects.
  • On a separate set of 127 moments where the human choice was already strong, model suggestions were preferred only about 20% of the time. The article does not give a model-by-model breakdown for that check.
  • The 90% practical ceiling uses an answer generated with knowledge of the full session. It is a hindsight reference, not an RSI threshold or an ordinary model result.
  • The source reports rounded percentages. Model-specific sample counts, confidence intervals, judge version and full task data are not supplied here.
Published values from the chart

Exact printed labels from the original chart, not estimates from bar lengths. Dates identify the models as labeled, not separate evaluation cohorts. All rows concern the same selected 129-session experiment. Ties are not counted as better; no unprinted remainder values have been calculated.

Choosing a better next step in a research problem
Model as labeledModel date on chartJudged betterTie
Claude Haiku 3March 202422%10%
Claude Sonnet 4May 202548%11%
Claude Sonnet 4.5September 202550%11%
Claude Haiku 4.5October 202545%11%
Claude Opus 4.5November 202551%10%
Claude Sonnet 4.6February 202645%13%
Claude Opus 4.6February 202655%14%
Claude Opus 4.7April 202659%12%
Claude Mythos PreviewApril 202664%9%

Related benchmarks and indicators

What to keep in mind

  • The 129 examples were selected because the human next step had room for improvement; this is not a representative comparison with researchers.
  • A separate Claude judged proposals using the eventual session outcome. Proposals were not tested by running complete alternative research projects.
  • On a separate set of 127 moments where the human choice was already strong, model suggestions were preferred only about 20% of the time. The article does not give a model-by-model breakdown for that check.