What the Productivity Studies Actually Found
The controlled experiments measuring AI's effect on work output, why their results range from large gains to measured slowdowns, and what distinguishes the settings.
There is now a body of randomised and quasi-experimental evidence on generative AI and work output. The results do not agree, and the disagreement is informative: the effect depends strongly on the task, the worker's prior skill, and whether the measured output is what the organisation actually values.
The headline results
Professional writing. In a randomised experiment on mid-level professional writing tasks, participants with access to a language model completed tasks in substantially less time, on the order of 40 percent, with graded output quality also improving, and the gains were concentrated among lower-performing participants so that inequality between workers narrowed (Noy and Zhang, 2023, Science 381:187-192).
Customer support. A field study of a large deployment to customer support agents found an increase in issues resolved per hour of roughly 14 percent on average, with the effect strongly skewed by experience: novice and low-skilled workers improved substantially, around a third, while the most experienced agents saw little or no gain (Brynjolfsson, Li and Raymond, Generative AI at Work, QJE).
Programming, controlled task. A randomised trial on a self-contained task, implementing an HTTP server in JavaScript, found developers with GitHub Copilot completed it around 56 percent faster (Peng et al., 2023, arXiv:2302.06590).
Programming, real repositories. A randomised trial with 16 experienced open-source developers on 246 tasks in mature repositories they averaged five years of experience with found the opposite: allowing AI tools increased completion time by 19 percent. The developers predicted a 24 percent speedup beforehand and estimated a 20 percent speedup afterwards, having in fact been slower (METR, 2025, arXiv:2507.09089).
Reconciling them
The two programming results are not contradictory once the settings are compared. A self-contained greenfield task with a clear specification is where a model contributes most: the developer has no context advantage, the code is standard, and verification is mechanical. A mature repository the developer knows deeply is where the model contributes least: the human's context is the scarce input, and the model's suggestions require review against knowledge it does not have.
The pattern across all four studies is consistent. Gains are largest where the task is well specified, the output is verifiable, and the worker's prior expertise is low. They shrink or reverse where the task requires deep context, where verification is expensive, and where the worker is already expert.
The METR perception gap is the finding with the widest implication. Participants believed they had been sped up while having been slowed, which means self-reported productivity is not evidence, and most organisational claims about AI productivity are self-reported.
When it breaks
Measured output is not always value. Faster resolution of support tickets, more code committed and more words written are measurable and are proxies. Whether the underlying outcome improved, the customer's problem solved, the software maintainable, the document useful, is a separate and harder measurement.
Novelty and learning both operate. Early measurements include the cost of learning the tool and the enthusiasm of trying it, pulling in opposite directions, so short studies are unreliable in both directions.
Task selection determines the result. A study can produce almost any effect size by choosing the task, and comparing headline numbers across studies without comparing settings is the error that makes this literature look chaotic.
Individual gains do not aggregate simply. If a bottleneck lies elsewhere, speeding up one step changes nothing at the system level. Firm-level and economy-level productivity effects have so far been much harder to detect than task-level ones, which is what the productivity paradox predicts.
12 flashcards for this concept
Click a card to reveal the answer.