operational metric

Experienced developer productivity

Randomized access to AI tools on real repository issues.

Observed R&D automation

What this measure tells us

Operational counterpoint to benchmark performance.

  • Experienced developers in familiar open-source repositories; not representative of all work.
  • Later study reports selection bias, so no unqualified trend.

Results by version

Late-2025 study update

2026-02-24

Selection effects prevent a reliable current effect estimate.

Published results · Late-2025 study update
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—METRInterpretability assessmentMETR reports that selection effects make the later experiment unreliable for estimating current productivity gains.2026-02-24productivity-later-pindependent evaluation

Version lineage

Reference points

No applicable reference points are published for this measure.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation serviceunknown

Official sources