Directory

Benchmarks & indicators

Measures of AI research and self-improvement ability.

benchmark

CoBench

Diagnose historical internal R&D failures from infrastructure snapshots.

AI-R&D capability
Reported result55.8%2.1 · Reported scoreClaude Opus 5.5 · 2026-09-22Other settings are available.See all results
Estimated requirement85%
Owner / creator: Anthropic
benchmark

KernelGen 1P

Optimize correct kernels for OpenAI first-party hardware.

AI-system improvement
Reported result66.72%Mean kernel optimization rewardGPT-6 Astra · 2026-09-03See all results
Owner / creator: OpenAI
benchmark

NanoGPT

Reduce small-model training time to a target validation objective.

AI-system improvement
Reported resultResults published as a chartExact point values are not printed.GPT-6 Astra · 2026-09-03Evidence record · Original source ↗
Human reference0.7238normalized scoreVersion Astra card snapshot, September 2026
Owner / creator: OpenAI
benchmark

PostTrainBench Lite

Improve a pretrained model within a five-hour GPU budget.

AI-system improvement
Reported resultResults published as a chartExact point values are not printed.GPT-6 Astra · 2026-09-03Evidence record · Original source ↗
Baseline reference1normalized scoreVersion Astra card snapshot, September 2026
Owner / creator: OpenAI
benchmark

PaperBench

Replicate research papers against author-developed rubrics.

AI-R&D capability
Reported result21.0 ± 0.8%Standard errorOriginal paper v1 · Mean replication scoreClaude 3.5 Sonnet (New) / BasicAgent · 2025-04-02 · 12 hoursOther settings are available.See all results
Owner / creator: OpenAI
benchmark

MLE-bench

ML competition engineering under bounded compute.

AI-R&D capability
Reported result93.80%Revised, Astra card snapshot · Mean percentile against reference solutionsGPT-6 Astra · 2026-09-03 · Max reasoningSee all results
Owner / creator: OpenAI
benchmark

RE-Bench

Research engineering environments with expert baselines.

AI-R&D capability
Reported result1.27normalized scoreGoogle model-card assessment, February 2026 · Human-normalised average scoreGemini 3.1 Pro (Deep Think) · 2026-02-19See all results
Baseline reference1normalized score
Owner / creator: METR
benchmark

Task-completion time horizons

Human-task duration associated with 50% agent success.

Research autonomy
Reported resultaround 11.3 hoursConfidence interval: 95% CI: 5–40 hoursCheating counted as failure; METR considers this estimate not reliable.TH1.1, GPT-5.6 Sol June 2026 assessment · 50% task-completion horizonGPT-5.6 Sol / METR horizon evaluation · 2026-06-26See all results
Owner / creator: METR
benchmark

AI4AI-Bench

Rewrite training algorithms in frozen research repositories.

AI-system improvement
Reported result0.250normalized scoreMean normalized scoreClaude Opus 5 / Claude Code effort aggregate · 2026-08-20 · 4 hours explorationOther settings are available.See all results
Baseline reference0.1normalized score
Owner / creator: AI4AI-Bench authors
operational metric

Experienced developer productivity

Randomized access to AI tools on real repository issues.

Observed R&D automation
Reported result-18%Confidence interval: -38% to +9% · Source-reported confidence interval; level not specified in summaryLate-2025 study update · AI-assisted task-time changeMETR · 2026-02-24 · Returning original-study participantsOther settings are available.See all results
Also reportedMETR reports that selection effects make the later experiment unreliable for estimating current productivity gains. · 2026-02-24Evidence record · Original source ↗
Owner / creator: METR