Research engineering environments with expert baselines.
AI-R&D capability
Reported result1.27normalized scoreGoogle model-card assessment, February 2026 · Human-normalised average scoreGemini 3.1 Pro (Deep Think) · 2026-02-19See all results
Human-task duration associated with 50% agent success.
Research autonomy
Reported resultaround 11.3 hoursConfidence interval: 95% CI: 5–40 hoursCheating counted as failure; METR considers this estimate not reliable.TH1.1, GPT-5.6 Sol June 2026 assessment · 50% task-completion horizonGPT-5.6 Sol / METR horizon evaluation · 2026-06-26See all results
Randomized access to AI tools on real repository issues.
Observed R&D automation
Reported result-18%Confidence interval: -38% to +9% · Source-reported confidence interval; level not specified in summaryLate-2025 study update · AI-assisted task-time changeMETR · 2026-02-24 · Returning original-study participantsOther settings are available.See all results
Also reportedMETR reports that selection effects make the later experiment unreliable for estimating current productivity gains. · 2026-02-24Evidence record · Original source ↗