Measuring AI-Assisted Developer Productivity
Why credible studies of the same tools report everything from a 19 percent slowdown to a 56 percent speedup, what each study design can and cannot establish, and how to read an effect size before repeating it.
Four randomised controlled trials of AI coding tools, all competently run, report effects that span a factor of three and two signs. A timed lab task found the treated group finished 55.8 percent faster (Peng et al., 2023, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv:2302.06590). Three field experiments pooling 4,867 developers at Microsoft, Accenture and a Fortune 100 firm found 26.08 percent more completed tasks, with a standard error of 10.3 percent (Cui et al., The Effects of Generative AI on High-Skilled Work, Management Science, 2025). An enterprise trial with 96 Google engineers found roughly 21 percent less time on a complex task, with a wide interval (Paradis et al., 2024, arXiv:2410.12944). And sixteen experienced open-source maintainers working in their own mature repositories took 19 percent longer (Becker et al., 2025, arXiv:2507.09089). None of these is wrong. They measured different things.
The measures are not interchangeable
Three distinct quantities get reported under the single word "productivity", and the differences between them explain most of the spread.
Time on a defined task is what the lab and enterprise trials measure. It is the cleanest causal estimate and the narrowest: it covers authoring a known artefact and excludes deciding what to build, integrating it, and living with it.
Throughput of units of work is what field experiments and telemetry studies measure, usually as completed tasks, commits, or merged pull requests. An observational study of Microsoft's early-2026 rollout of command-line coding agents found adopters merged about 24 percent more pull requests than their own trajectory predicted, sustained over four months (Murphy-Hill et al., 2026, arXiv:2607.01418). A merged pull request is a countable proxy, not a unit of value, and the authors say so.
Delivery and quality outcomes are what the DORA programme measures: throughput, change failure rate, rework, time to restore. Its 2025 edition reports AI adoption now correlating positively with delivery throughput while also correlating with higher instability, and it tested whether fast repair offsets the instability and found that it does not (DORA, 2025, State of AI-Assisted Software Development).
A tool can move all three in different directions at once. Faster authoring with more rework and flat delivered value is a coherent outcome, not a contradiction.
The evidence hierarchy
Rank a claim by what its design can support, not by how large the number is.
- Randomised trial on real tasks in the subject's own codebase. Strongest causal claim; tiny samples; expensive. METR's trial is the only one in this class to date.
- Randomised trial on an assigned task. Clean causal claim about authoring; the task is chosen by the experimenter, which is where most of the generalisation risk sits.
- Field experiment with randomised access. Real work, real incentives, randomised treatment, but access is not use, so the estimate is diluted by non-adopters unless the paper models adoption.
- Quasi-experimental telemetry. Difference-in-differences on repository histories can cover thousands of projects. One such study of agent adoption in open source found speed gains concentrated in projects for which the agent was their first AI tool, with little or short-lived gain where an AI IDE assistant was already in use (Agarwal, He and Vasilescu, 2026, AI IDEs or Autonomous Agents?, MSR 2026, arXiv:2601.13597).
- Self-report. Useful for measuring beliefs, and nothing else. In METR's trial, participants forecast a 24 percent speedup, experienced a 19 percent slowdown, and afterwards still estimated they had been sped up by 20 percent.
Vendor telemetry reports sit outside this hierarchy because their samples, definitions and baselines are rarely reproducible. Read them for direction, not magnitude, and check whether a figure quoted as this year's was published last year.
Why the effect is a function, not a number
The moderators that the studies agree on are the ones worth carrying around. Gains fall as repository familiarity rises, because the model's advantage is largest exactly where the developer's own recall is weakest. Gains fall as the codebase's implicit constraints grow, since a mature project encodes conventions no prompt conveys. Gains fall as verification cost rises relative to authoring cost. And gains concentrate among less experienced developers, a finding reported independently by the Copilot trial and the three field experiments.
That is why "does AI make developers faster?" has no scalar answer, and why an organisation quoting a single percentage is almost certainly quoting the study whose population least resembles its own.
When it breaks
The control arm is eroding. METR's follow-up, which ran from August 2025 with 57 developers, 143 repositories and over 800 tasks, returned a point estimate of 18 percent faster for returning developers with an interval from 38 percent faster to 9 percent slower, and 4 percent faster for new recruits: both intervals include zero. METR's own account of why it changed the design is the important part. Between 30 and 50 percent of developers said they were declining to submit tasks they did not want to do without AI, and some were reluctant to be randomised into the no-AI arm at all (METR, 2026, We are Changing our Developer Productivity Experiment Design). Selection on the task list breaks the randomisation that the design depends on.
Snapshot results age in months. Every estimate above is bound to a tool generation. The METR trial used Cursor Pro with Claude 3.5 and 3.7 Sonnet in early 2025; the Microsoft study measured command-line agents in 2026. Quoting the first as evidence about the second is a category error, in either direction.
Counting artefacts rewards producing artefacts. Pull requests, commits and lines are cheap to inflate once generation is cheap. A measurement programme built on them will record a large improvement during exactly the period when review queues and change failure rate are deteriorating. Pair any throughput measure with a stability measure or do not run the programme.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Peng et al., 2023, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv:2302.06590 arxiv.org
- Cui et al., The Effects of Generative AI on High-Skilled Work, Management Science, 2025 papers.ssrn.com
- Paradis et al., 2024, arXiv:2410.12944 arxiv.org
- Becker et al., 2025, arXiv:2507.09089 arxiv.org
- Murphy-Hill et al., 2026, arXiv:2607.01418 arxiv.org
- DORA, 2025, State of AI-Assisted Software Development services.google.com
- Agarwal, He and Vasilescu, 2026, AI IDEs or Autonomous Agents?, MSR 2026, arXiv:2601.13597 arxiv.org
- METR, 2026, We are Changing our Developer Productivity Experiment Design metr.org
6 flashcards for this concept
Click a card to reveal the answer.