CoBench
Benchmark by Anthropic
2.1
Diagnose historical internal R&D failures from infrastructure snapshots.
55.8%
Claude Opus 5.5
RSI Progress
What has been measured, what changed, and what the evidence tells us about recursive self-improvement.
Selected site evidence · Summary reviewed
Where are we? ·
Assessment from the paper by Duan et al. (Sep. 2026) ↗
Levels describe different kinds of autonomy, not necessarily a required sequence.
The top rung, L5: AI improves how it improves
No single benchmark or percentage captures progress toward RSI.
Each score belongs to one system, test version, and setup.
Doing research tasks well is not the same as controlling the improvement loop.
Benchmark by Anthropic
2.1
Diagnose historical internal R&D failures from infrastructure snapshots.
55.8%
Claude Opus 5.5
Benchmark by OpenAI
Astra card snapshot, September 2026
Resolve real bugs in internal research experiments.
78.05%
GPT-6 Astra
Benchmark by OpenAI
Astra card snapshot, September 2026
Optimize correct kernels for OpenAI first-party hardware.
66.72%
GPT-6 Astra
Research, debugging, and engineering work involved in building AI.
6 measuresChanges to training, post-training, algorithms, data, and tools.
5 measuresHow independently and reliably systems complete extended work.
1 measureReported contribution in real organizational work.
3 measuresEvidence that improved systems contribute to later improvements.
Explore evidenceControls for contamination, gaming, overfitting, and transfer.
Explore evidence