Evidence record
Experienced developer productivity
Late-2025 study update · METR · 2026-02-24
Reported result · qualitative
METR reports that selection effects make the later experiment unreliable for estimating current productivity gains.
Interpretability assessment · qualitative
independent evaluation · extraction review: agent checked
Source and extraction
Published 2026-02-24
- We are Changing our Developer Productivity Experiment DesignIntroduction and selection-effects discussionOriginal source ↗
Evaluation setup
| selection rule | Volunteer participation and self-selected submitted tasks |
|---|---|
| comparability caveats | Selection and concurrent-agent time measurement biases. |
Not reported: split, task count, task snapshot, attempts per task, run count, aggregate method, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Reported raw effects are not used as a comparable trend.
- Experienced developers in familiar open-source repositories; not representative of all work.
- Later study reports selection bias, so no unqualified trend.