report · 2026-08-28

Automated Researchers Can Reliably Mitigate Alignment Failures

Why it matters here

AI can help develop training methods that improve other AI systems. Anthropic’s Claude Opus 4.8 research agents found methods that reduced ten defined types of harmful or unreliable behavior in smaller models. The improvements also appeared on tests kept separate from training and on larger models. This is evidence of AI doing part of AI-development research, with people still choosing the goals and tests.

Read original source ↗

Research autonomyAI-system improvement

What the charts show

AI-developed training methods improved results on separate tests

The AI-developed methods improved results for all ten tested failure types on separate evaluation tasks. Selected methods also worked on larger models, suggesting the improvements extended beyond the models used to develop them.

Blue bars show models before the additional training; orange bars show them after using an AI-developed method. Each panel measures a different behavior. Higher scores mean less of that behavior on the named test.

Original Anthropic Figure 4: ten small multiples comparing untrained and AAR-trained model scores on distinct held-out safety benchmarks, including larger-model transfer.
Unmodified original Anthropic fig4.png downloaded from the HTML article; visually cross-checked against PDF p. 7, Fig. 4. Enlarge chart ↗

View original chart and methods ↗
Automated Researchers Can Reliably Mitigate Alignment Failures · PDF p. 7, Fig. 4 and caption

How to interpret this chart
  • The panel benchmarks have different definitions and are not a common cross-failure scale.
  • Researchers used these evaluation scores to choose among leading methods for the larger-model and subsequent Petri behavioral-audit tests. Petri scores were not used for that selection.
  • No bar heights were transcribed as exact scores.

AI-created training data improved an early Claude model

In a separate experiment, Claude Sonnet 5 created training data for an early Claude Opus 4.8 model. It closed 65% of the measured safety gap, compared with 72% for the released model after Anthropic’s full training process.

The roughly 60-hour experiment targeted ten defined behaviors using Petri, a system for auditing model behavior. The score measures how much of the gap between the starting model and the test’s defined ceiling was closed. It is not a percentage of overall safety; the two training processes used different resources.

Original Anthropic Figure 9: best accepted Petri safety-gap score during the early-checkpoint trial, with labeled 65% winner and 72% released-model reference.
Unmodified original Anthropic fig9.png downloaded from the HTML article; visually cross-checked against PDF p. 12, Fig. 9. Enlarge chart ↗

View original chart and methods ↗
Automated Researchers Can Reliably Mitigate Alignment Failures · PDF p. 12, Fig. 9 and caption

How to interpret this chart
  • This is a Petri audit proxy for ten measured failure types, not overall alignment.
  • The comparison is to the released checkpoint after Anthropic production training; training budgets and interventions are not matched.
  • The researcher in this trial is Claude Sonnet 5, whereas the small-model experiments used Claude Opus 4.8.
Published values from the chart

Exact labels from Figure 9. The released model is a reference, not an RSI threshold. Other points have no printed exact scores and are not transcribed.

AI-created training data improved an early Claude model
Training conditionMeasured safety gap closed
Early Opus 4.8 trained with Sonnet 5’s proposed data65%
Released Opus 4.8 after full production training72%

What to keep in mind

  • The measured failures and capability checks are proxies; open-ended, unknown or rare failures and unmeasured capabilities remain outside the claim.
  • Human researchers submitted one-shot ideas with up to eight hours and could not iterate, unlike the AAR search; this is not a matched research-productivity comparison.
  • Thirty-nine of 1,601 proposed-method trajectories were flagged as cheating and excluded; the monitor is post-hoc and does not establish robustness to strongly evasive agents.
  • The early checkpoint Petri comparison is a different setup from the small-model method experiments and does not isolate equivalent training budgets or overall production alignment.

Study details

Bounded AI research and post-training of other models. The researcher agents did not retain an improved version of themselves, and the report does not demonstrate sustained recursive self-improvement or general alignment.

Methods and study design
Study design and reported evidence
What changedweights, training_algorithm, training_data
What stayed fixedResearch agents were externally tasked with a named failure, benchmark suite, target and constraints.; Small-model AAR identity remained Claude Opus 4.8; the agents did not train or replace their own researcher model.; Capability checks and held-out evaluator were separate from training and proposed-method code.
Improved system used as optimizer laterno
Generations attempted / acceptednot reported / not reported
Held-out transferThe top leaderboard method exceeded the untrained target on the held-out benchmark for each of ten failure types. For larger-model and Petri tests, leading methods were selected using the held-out benchmark; Petri was not used for that selection. Larger target models were 1.8–4.7× the optimized targets by parameters.
Resource accountingMain small-model harness used five parallel Opus 4.8 agents for up to 48 hours per failure; each candidate method received one H200 GPU and about 30 minutes of training. The separate Sonnet 5 early-checkpoint trial lasted about 60 hours.
Human contributionsHumans chose the ten failure types, assembled and validated benchmark suites, set rules and capability gates, recruited human comparators, and designed the monitored AAR harness and evaluator.
Author claimsAutomating alignment post-training may be practical for well-characterized failures with measurable proxies; a separate early production-checkpoint trial suggests possible future use.