Bounded prototypes and emerging research or industrial systems.
Where are we? · autonomy evidence
RSI autonomy frontier
Explore which parts of improving an AI system the AI can control, and which decisions humans still make. This page explains the levels of independence and shows the evidence for each across research areas.
Current evidence frontier
Inspect source records →Lower and intermediate levels have broad evidence; L3–L4 are more domain-dependent.
Some bounded meta-improvement and transfer; reliable multi-generation accumulation remains unresolved.
Not established by this survey’s reviewed evidence under comparable resources.
Coverage is insufficient for an independent broad or global stage assignment.
Bounded mechanism reuse in the selected AIDE² experiment.
Partial evidence of better research harnesses; outer-improver advantage inconclusive.
Not established in this evidence set; not a claim that no other evidence exists.
Evidence summary
RSI Progress assessment ·Bounded recursive demonstrations
Robust full-loop RSI is not established in the reviewed evidence.
AI research capability
Measured task gains: Research and engineering gains appear in selected evaluations.
Selected evaluations measure research and engineering tasks. AIDE² reports improved research harness performance and transfer to held-out benchmarks. Task performance alone does not show autonomous control of an improvement loop.
Mechanism records: AIDE2
Control of the improvement loop
Partial control: People still set key goals, evaluations and resource limits.
Control depends on the system and domain. The Duan survey distinguishes intervention choice, learning experience, deployment adaptation and future-improver revision. Human-set objectives, evaluation, resources and release decisions remain external controls in the selected experiments.
Structural inheritance
Bounded demonstrations: Modified mechanisms are retained and reused in DGM and AIDE².
DGM and AIDE² provide bounded examples of modified improvement mechanisms being retained and used again. DGM changes agent scaffolds over fixed foundation models. AIDE² separates inner harness search from outer-improver revision. These are distinct experimental systems.
Better future improvement
Inconclusive: Better task performance has not settled whether the improver improves.
An improved task solver does not automatically make a better improver. AIDE² reports harness gains, but its evolved outer-improver comparison is inconclusive. The reviewed DGM evidence does not settle robust meta-improvement under a defensible matched comparison.
Sustained gains and acceleration
Not established: Reliable multi-generation gains and acceleration remain unproven here.
Search iterations, accepted updates and inherited generations are different quantities. These selected studies do not establish reliable, sustained improvement of the improver or acceleration under comparable resources. This is a limit of the reviewed evidence, not an exhaustive absence claim.
Limits of this assessment
- Selected public evidence, not an exhaustive or continuously monitored literature review.
- Different systems and benchmarks cannot be assembled into evidence of a single integrated RSI loop.
- Source extraction has been agent-checked; no independent experimental replication or human review is claimed.
- This qualitative summary is not a percentage of completion, probability of RSI or time-to-RSI forecast.
B0–L5 autonomy ladder
Full definitions and criteria →- B0Refine outputIn-task AI improvement
- L1Execute improvementsImprovement execution autonomy
- L2Choose how to improveImprovement strategy autonomy
- L3Choose what to learnExperience-acquisition autonomy
- L4Adapt from deploymentEnvironment adaptation autonomy
- L5Improve the improvement mechanismRecursive inheritance autonomy
Evidence by research area
Evidenced mechanisms: L2
Strong L2; early L3. Survey examples include CASCADE and CORAL; these examples are survey coverage, not independently audited system records.
External controls: Scientific objectives; Validation, reproducibility and safety constraints
Evidenced mechanisms: L2 · L3
L2–L3 main frontier; simulated L4 evidence. Survey examples include Voyager, EnvGen and POET; physical-world integration remains limited.
External controls: Environment interfaces; Evaluators; Physical safety boundaries
Evidenced mechanisms: L2
Mature L2; emerging L3; bounded structural L5. The survey cites Self-Harness, learner-conditioned SWE systems, SICA and DGM; L4 deployment adaptation remains largely absent.
External controls: Objectives and specifications; Benchmark tests; Archive and parent-selection machinery
Evidenced mechanisms: L2
Mature L2; limited L3-like evidence; simulated L4. Survey examples include EvoClinician and EvoPatient; this is not evidence of autonomous clinical deployment.
External controls: Patient populations and task distribution; Clinical evaluation; Deployment and safety gates
Evidenced mechanisms: L5
Bounded structural L5 in AIDE²; no representative domain-wide frontier assigned. Capability-only lab benchmarks do not establish autonomous inheritance.
External controls: Protected evaluation; Task families; Model substrate; Cost and release authority
Evidence records
Science
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L2
Strong L2; early L3. Survey examples include CASCADE and CORAL; these examples are survey coverage, not independently audited system records.
External controls: Scientific objectives; Validation, reproducibility and safety constraints
Science
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L3
Strong L2; early L3. Survey examples include CASCADE and CORAL; these examples are survey coverage, not independently audited system records.
External controls: Scientific objectives; Validation, reproducibility and safety constraints
Embodied intelligence
source author assessment · assessed 2026-09-10 by Duan et al.
L2–L3 main frontier; simulated L4 evidence. Survey examples include Voyager, EnvGen and POET; physical-world integration remains limited.
External controls: Environment interfaces; Evaluators; Physical safety boundaries
Embodied intelligence
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L4
L2–L3 main frontier; simulated L4 evidence. Survey examples include Voyager, EnvGen and POET; physical-world integration remains limited.
External controls: Environment interfaces; Evaluators; Physical safety boundaries
Software engineering
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L2
Mature L2; emerging L3; bounded structural L5. The survey cites Self-Harness, learner-conditioned SWE systems, SICA and DGM; L4 deployment adaptation remains largely absent.
External controls: Objectives and specifications; Benchmark tests; Archive and parent-selection machinery
Software engineering
source author assessment · assessed 2026-09-10 by Duan et al.
Mature L2; emerging L3; bounded structural L5. The survey cites Self-Harness, learner-conditioned SWE systems, SICA and DGM; L4 deployment adaptation remains largely absent.
External controls: Objectives and specifications; Benchmark tests; Archive and parent-selection machinery
Healthcare
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L2
Mature L2; limited L3-like evidence; simulated L4. Survey examples include EvoClinician and EvoPatient; this is not evidence of autonomous clinical deployment.
External controls: Patient populations and task distribution; Clinical evaluation; Deployment and safety gates
Healthcare
source author assessment · assessed 2026-09-10 by Duan et al.
Mature L2; limited L3-like evidence; simulated L4. Survey examples include EvoClinician and EvoPatient; this is not evidence of autonomous clinical deployment.
External controls: Patient populations and task distribution; Clinical evaluation; Deployment and safety gates
AlphaEvolve / Gemini ensemble
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: B0
Duan et al. place the task-specific AlphaEvolve program-search boundary at B0 in their industry table. Deployment of generated artifacts is a different boundary from autonomous inheritance by the improver.
External controls: Task-specific evaluators; Human production integration; Human contributions: Experts define evaluators and integrate validated changes.
Supporting records: alphaevolve-study
Science
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L1
Survey describes persistent model and memory updates as L1 examples; this does not mean all science systems share that stage.
External controls: Training and selection procedure
Darwin Gödel Machine / v1
source author assessment · assessed 2026-09-10 by Duan et al.
Framework levels: L5
The survey describes bounded L5 characteristics in DGM while keeping archive management and parent selection outside self-modification.
External controls: Archive selection; Benchmark; Base model; Human contributions: Human-designed evaluation, initial agent, safety boundaries and experiment setup.
Supporting records: dgm-study
AIDE² (September 2026 study)
tracker evidence synthesis · assessed 2026-09-26 by RSI Tracker (Codex evidence synthesis)
Framework levels: L5
Revised research harness is retained and runs later; the separate outer-improver test invokes an evolved harness. Stronger effective recursion remains inconclusive.
External controls: Private evaluator; Task families; Budget; Outer selection rule; Human contributions: Evaluator design, seed agents and resource limits.
Supporting records: aide2-study
Contrary records: rec-aide2
AIDE² (September 2026 study)
tracker evidence synthesis · assessed 2026-09-26 by RSI Tracker (Codex evidence synthesis)
Framework levels: L5
Same harness mechanism viewed at the software-engineering boundary; not a domain-wide stage assignment.
External controls: Task families; Protected grading; Human contributions: Evaluator design, seed agents and resource limits.
Supporting records: aide2-study
AIDE²: inherited harness and outer-improver test
Evidence synthesis by RSI Tracker (Codex evidence synthesis)
Structural L5: demonstrated · created yes, retained yes, inherited yes, invoked yes
Effective L5: partial · budget matched yes · independent evaluation no
Recursive acceleration: not demonstrated
Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.
Darwin Gödel Machine: scaffold-level recursive improvement
Evidence synthesis by RSI Tracker (Codex evidence synthesis)
Structural L5: demonstrated · created yes, retained yes, inherited yes, invoked yes
Effective L5: unclear · budget matched unknown · independent evaluation unknown
Recursive acceleration: not demonstrated
Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.
AlphaEvolve: iterative artifact optimization
Evidence synthesis by RSI Tracker (Codex evidence synthesis)
Structural L5: not demonstrated · created unknown, retained unknown, inherited unknown, invoked unknown
Effective L5: unclear · budget matched unknown · independent evaluation unknown
Recursive acceleration: not demonstrated
Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.
No records match these filters.
Frontier evidence timeline
AIDE² (September 2026 study) · AI R&D / frontier model development
Supports: Technical report documents inherited harnesses and the outer-improver test.
Does not establish: No decisive advantage in that test; acceleration remains unestablished.
Domain evidence · Science
Supports: A source-attributed taxonomy and domain synthesis; an evidence-interpretation event.
Does not establish: No measured global-stage transition.
Darwin Gödel Machine / v1 · Software engineering
Supports: Reusable modified agents generate descendants in bounded coding search.
Does not establish: No matched proof of improving the improvement rate.
No events match these filters.