Evidence record
Experienced developer productivity
Late-2025 study update · METR · 2026-02-24
Correction or update linkedIndexed late-2025 raw cohort estimates, preserving METR’s selection-bias warning.
Reported result · numeric
-18%
Confidence interval: -38% to +9% · Source-reported confidence interval; level not specified in summary
Returning original-study participants
AI-assisted task-time change · percent
independent evaluation · extraction review: agent checked
Source and extraction
Published 2026-02-24
- We are Changing our Developer Productivity Experiment DesignIntroduction, Raw results paragraph (February 24, 2026)Original source ↗
Evaluation setup
| split | Returning original-study participants |
|---|---|
| selection rule | Volunteer participation and self-selected submitted tasks |
| aggregate method | Raw cohort-specific AI-assisted task-time change |
| comparability caveats | Severe task and participant selection effects; concurrent-agent timing may be unreliable. Not a representative current productivity effect. |
Not reported: task count, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Raw estimate, not a reliable causal estimate of current developer productivity.
- Study participation and submitted tasks were selected; some concurrent-agent task times are unreliable.
- Experienced developers in familiar open-source repositories; not representative of all work.
- Later study reports selection bias, so no unqualified trend.