RSI Progress

Tracking AI’s progress toward recursive self-improvement

What has been measured, what changed, and what the evidence tells us about recursive self-improvement.

Measures
14
Public results
79
Papers & reports
16

Progress toward self-improving AI

Selected site evidence · Summary reviewed

  • AI is taking on useful parts of AI development. Systems can investigate research problems, run experiments, improve training methods and optimize code. Some improvements have already been used in real model training.
  • Progress is uneven. The reports show gains on research tasks and a growing ability to handle longer tasks, but also experiments that stall or fail. Stronger benchmark scores do not automatically translate into faster real-world work.
  • People still guide the process. AI can lead parts of research, while people set goals, evaluation standards and resource limits. The amount of independent control varies across systems and research areas.
  • Sustained self-improvement remains unresolved. Some agents retain useful changes and reuse modified versions of themselves. The tracked evidence has not established a reliable process in which successive generations also become better at finding and producing further improvements.

Where are we? ·

How much of its own improvement does AI control?

Assessment from the paper by Duan et al. (Sep. 2026) ↗

ScienceDemonstratedDemonstratedDemonstratedEarly evidenceNot establishedNot established
Embodied intelligenceDemonstratedDemonstratedDemonstratedDemonstratedEarly evidenceNot established
Software engineeringDemonstratedDemonstratedDemonstratedEarly evidenceNot establishedEarly evidence
HealthcareDemonstratedDemonstratedDemonstratedEarly evidenceEarly evidenceNot established
DemonstratedEarly evidenceNot established

Levels describe different kinds of autonomy, not necessarily a required sequence.

What each level means
B0 Refine output
Changes refine an output during the current task or session. Future independent tasks inherit no accepted system update.
L1 Execute improvements
Humans prescribe the target, update procedure and acceptance criteria. AI executes the procedure and accepted changes persist into later tasks or rounds.
L2 Choose how to improve
AI diagnoses weaknesses and chooses interventions and experiments. Humans continue to set objectives, task boundaries and evaluation criteria.
L3 Choose what to learn
AI chooses or generates subsequent learning experience based on the evolving learner’s state. What it learns from changes as the learner changes.
L4 Adapt from deployment
Ongoing deployment or environmental interaction determines persistent changes in memory, skills, code, harnesses or parameters reused on later operational tasks.
L5 Improve the improvement mechanism
The procedure governing future improvement is itself revised, retained and invoked in subsequent improvement rounds. It may be an improver, search or research policy, evaluator or successor generator.

The top rung, L5: AI improves how it improves

Structural L5Bounded prototypes and emerging research or industrial systems.
Effective L5Some bounded meta-improvement and transfer; reliable multi-generation accumulation remains unresolved.
Recursive accelerationNot established by this survey’s reviewed evidence under comparable resources.

Read the evidence in context

How we interpret results →
  1. No single score

    No single benchmark or percentage captures progress toward RSI.

  2. Results keep their context

    Each score belongs to one system, test version, and setup.

  3. Capability ≠ autonomy

    Doing research tasks well is not the same as controlling the improvement loop.

Selected evidence

Browse all measures

Evidence categories

Recent evidence activity

Indexed: Claude Opus 5.5 System Card

Indexed: Measurements for understanding the pace of AI development inside frontier labs

Added the source-attributed B0–L5 framework and historical domain assessment.

Indexed: GPT-6 Astra System Card

Show 20 earlier updates

Indexed four exact comparator labels from the Astra-card debugging figure; this is historical evidence, not a new model release.

Indexed: AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Indexed the paper’s separate Claude Opus 5 medium-effort result, distinct from its across-effort mean.

Indexed: Gemini 3.7 Flash Frontier Safety Framework Report

Indexed METR’s conditional, non-robust GPT-5.6 Sol time-horizon estimate with its uncertainty.

Indexed: We are Changing our Developer Productivity Experiment Design

Indexed late-2025 raw cohort estimates, preserving METR’s selection-bias warning.

Added Google’s February 2026 numerical RE-Bench assessment.

Indexed: Time Horizon 1.1

Completed the printed TH1/TH1.1 comparison-table rows omitted from the tracker’s historical snapshot.

Indexed: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Indexed: Recent Frontier Models Are Reward Hacking

Indexed: Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Indexed: AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms

Indexed: PaperBench: Evaluating AI’s Ability to Replicate AI Research

Indexed omitted original PaperBench and distinct Code-Dev results from the 2025 paper.

Indexed the paper’s three-paper agent result and human best-of-three reference with different time accounting; no upstream benchmark change.

Indexed: Evaluating frontier AI R&D capabilities of language model agents against human experts

Indexed: MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Indexed original MLE-bench GPT-4o scaffold comparisons from the 2024 paper.

Explore the tracker