Evidence record
Experienced developer productivity
Early-2025 randomized trial · METR · 2025-07-10
Reported result · numeric
19% longer
Change in completion time · percent
independent evaluation · extraction review: agent checked
Source and extraction
Published 2025-07-10
Evaluation setup
| task count | 246 |
|---|---|
| selection rule | Random assignment of issues to AI allowed or disallowed |
| aggregate method | Estimated change in completion time |
| tool access | Chosen tools, primarily Cursor with Claude 3.5/3.7 Sonnet |
Not reported: split, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Positive means slower. Not an estimate for all developers or current tools.
- Experienced developers in familiar open-source repositories; not representative of all work.
- Later study reports selection bias, so no unqualified trend.