Reasoning & Evaluation

The Control Group Is Disappearing: Measuring AI-Assisted Development

Four credible randomised trials of AI coding tools report effects from a 19 percent slowdown to a 56 percent speedup. The newest problem is worse than the spread: developers now refuse to be randomised into the no-AI arm, and 30 to 50 percent withhold tasks they will not do without it. The counterfactual is eroding faster than the measurements are improving.

In February 2026, METR published an account of an experiment that did not work. The organisation had run the most-cited randomised trial in this area, in which sixteen experienced open-source maintainers took 19 percent longer to complete real tasks in their own repositories when AI tools were allowed (Becker et al., 2025, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, arXiv:2507.09089). The follow-up, running from August 2025 across 57 developers, 143 repositories and more than 800 tasks, was meant to settle whether that result held as the tools improved. It returned a point estimate of 18 percent faster for returning developers, with a confidence interval from 38 percent faster to 9 percent slower, and 4 percent faster for new recruits. Both intervals contain zero. The reason METR gave for redesigning the study is the part worth reading twice: surveyed developers said, at rates between 30 and 50 percent, that they were choosing not to submit tasks because they did not want to do those tasks without AI, and some were reluctant to be randomised into the no-AI condition at all (METR, 2026, We are Changing our Developer Productivity Experiment Design).

That is not a measurement problem of the ordinary kind. A trial that cannot populate its control arm with representative work has lost the comparison it was built on. Two and a half years after the first Copilot trial, the question "how much faster does AI make developers?" has become harder to answer than it was in 2023, and not because the evidence got thinner. It got thicker, better-designed, and mutually contradictory, while the behaviour it was trying to measure changed underneath it.

Why this matters: Every organisation now buying coding assistants and agents at four to five figures per engineer per year is making a capital allocation decision against evidence whose effect sizes span two signs and a factor of three. The studies do not disagree because some are wrong. They disagree because "productivity" names three different quantities, and AI moves them in different directions at once.

TL;DR

  • Four well-run randomised trials report: 55.8 percent faster on a timed greenfield task (Peng et al., 2023), 26.08 percent more completed tasks across 4,867 developers (Cui et al., Management Science, 2025), about 21 percent less time on a complex enterprise task (Paradis et al., 2024), and 19 percent slower on self-chosen tasks in mature repositories (Becker et al., 2025). All four can be right.
  • The spread is explained by the outcome variable, not by tool quality: time-on-assigned-task, throughput of artefacts, and delivered outcomes are three different measurements.
  • Generation got cheap; verification did not. Net speedup follows an Amdahl identity in which authoring is often only 25 to 40 percent of a task, so the ceiling on a generation-only improvement is low and the floor is negative.
  • Throughput gains land on a review queue whose capacity did not change. A team moving from 70 to 91 percent review utilisation sees mean queue wait rise roughly fivefold, with no change in reviewer behaviour.
  • DORA's 2025 survey of an industry at 95 percent adoption found throughput and instability both rising with AI use, and tested whether faster repair offsets the instability. It does not (DORA, 2025).
  • Self-report is now provably unreliable in this domain: METR's participants forecast a 24 percent speedup, measured a 19 percent slowdown, and afterwards still estimated a 20 percent speedup.
  • The strongest documented industrial gains are not in feature work but in mechanical migration, where a compiler and a test suite replace human judgement as the oracle: Google reports 74.45 percent of migration changes model-generated and roughly 50 percent total time saved (Ziftci et al., 2025).

At a Glance

The whole argument fits in one causal chain. Generation cost collapsed, which accelerates exactly one phase of software work; the phases downstream of it did not change capacity, so the saving is partly consumed by verification and partly converted into queue.

flowchart LR
  Gen["Generation cost collapses"] --> Author["Authoring phase accelerates"]
  Author --> Verify["Verification unchanged or slower"]
  Author --> Queue["More changes arrive at review"]
  Verify --> Net["Net task time: ambiguous sign"]
  Queue --> Wait["Queue wait rises non-linearly"]
  Wait --> Inst["Instability and rework rise"]
  Net --> Measure["What a trial measures"]
  Inst --> Measure
  classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
  classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
  classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
  classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
  classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
  class Gen blue
  class Author purple
  class Verify,Queue amber
  class Wait,Inst rose
  class Net,Measure slate

[IMAGE: Three small multiples side by side on a shared x-axis of study publication date, y-axis percentage effect, showing time-on-task studies, artefact-throughput studies, and delivery-stability studies as separate panels with their confidence intervals. Caption: "Same tools, three outcome variables, three different stories."]

Before the Question Got Hard

GitHub announced Copilot as a technical preview on 29 June 2021 and took it out of preview on 21 June 2022, which is the moment autocomplete-scale code generation became a line item rather than a demo. The first serious causal estimate followed quickly: a timed experiment in which developers implemented an HTTP server in JavaScript found the treated group finished 55.8 percent faster, with a 95 percent interval from 21 to 89 percent (Peng et al., 2023, The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv:2302.06590). That number travelled further than any other figure in the field, usually without its task description attached.

Field evidence arrived next, and it was better evidence. Three randomised experiments run inside the normal operations of Microsoft, Accenture and a Fortune 100 company, pooling 4,867 developers, estimated a 26.08 percent increase in completed tasks with a standard error of 10.3 percent, alongside 13.55 percent more weekly commits and 38.38 percent more weekly builds (Cui et al., The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers, Management Science, 2025). Google's own trial with 96 engineers on a complex enterprise task found roughly 21 percent less time spent, with the authors flagging a wide interval and explicitly warning against generalising from internal tooling in mid-2024 (Paradis et al., 2024, How much does AI impact development speed?, arXiv:2410.12944).

By mid-2025 the picture was a consensus in search of a counterexample, and METR supplied one. Its design was the strongest in the literature on the dimension that matters most for external validity: real tasks, chosen by the developers, in repositories they already maintained, with screen recordings to verify self-reported times. The result went the other way, and the authors were careful about scope, attributing the slowdown in part to high repository familiarity and to the size and maturity of the projects involved.

Then the measurement itself started to deform. DORA's 2025 survey described a population at 95 percent adoption spending around two hours a day with these tools. A difference-in-differences study of coding-agent adoption across open-source repositories found velocity gains concentrated almost entirely in projects for which the agent was their first AI tool (Agarwal, He and Vasilescu, 2026, AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development, MSR 2026, arXiv:2601.13597). METR redesigned its experiment. And the most recent large study abandoned randomisation altogether in favour of observing a real rollout: across tens of thousands of Microsoft engineers, adopters of command-line coding agents merged about 24 percent more pull requests than their own prior trajectory predicted, an effect that held across a four-month window (Murphy-Hill, Butler and Savelieva, 2026, Adoption and Impact of Command-Line AI Coding Agents, arXiv:2607.01418).

timeline
    title Five years of trying to measure this
    2021 : GitHub Copilot enters technical preview in June
    2022 : Copilot leaves preview as a paid product in June
    2023 : Peng et al. report 55.8 percent faster on a timed task
    2024 : Cui et al. field experiments find 26 percent more tasks completed
         : Paradis et al. find about 21 percent less time at Google
    2025 : Google publishes two industrial migration programmes
         : METR finds experienced maintainers 19 percent slower
         : DORA reports throughput and instability rising together
    2026 : CMU difference-in-differences finds gains only on first AI adoption
         : METR redesigns its trial after control-arm selection effects
         : Microsoft observes 24 percent more merged pull requests from agent adopters

[IMAGE: Annotated timeline strip with each year's study rendered as a small forest-plot marker showing point estimate and interval, so the widening disagreement after 2024 is visible as overlapping and opposite-signed intervals. Caption: "The evidence did not converge; it fanned out."]

What These Studies Are Actually Measuring

Three quantities, one word

Almost all of the apparent contradiction dissolves once the outcome variables are separated.

Time on a defined task is the cleanest causal estimate available and the narrowest in scope. It covers authoring a known artefact. It excludes deciding what to build, integrating it, reviewing it, and operating it. The 55.8 percent and the 21 percent figures live here.

Throughput of countable artefacts is what field experiments and telemetry measure: completed tasks, commits, merged pull requests. This is closer to what a business cares about and much further from a causal claim, because the counting unit is under the tool's influence. A merged pull request is a proxy for output; the Microsoft authors say exactly that about their own 24 percent.

Delivery and quality outcomes are the DORA quantities: throughput, change failure rate, rework, time to restore. The 2025 edition reports AI adoption correlating positively with delivery throughput while simultaneously correlating with higher instability, and it is the only study here to test the obvious optimistic hypothesis, that teams shipping faster also repair faster and end up even. The analysis found no support for that (DORA, 2025).

A tool that makes authoring four times faster, review 20 percent slower, and change failure rate 15 percent worse will produce a glowing report on measure two, an ambiguous report on measure one, and a troubling report on measure three. That is one coherent system, not three contradictory findings.

The identity that sets the ceiling

Write a task as phases with baseline time fractions \(f_i\) and per-phase speedup factors \(s_i\). Total time relative to baseline is

\[T = \sum_i \frac{f_i}{s_i}, \qquad \sum_i f_i = 1\]

For a four-phase decomposition of orientation, authoring, verification and rework, generation acts on \(f_{\text{author}}\) alone. If authoring is 30 percent of the task, then even \(s_{\text{author}} \to \infty\) caps the saving at 30 percent, and in practice \(s_{\text{author}}\) of 2 to 4 yields 15 to 22 percent. That is the entire upside of a generation-only improvement, and it explains why honest field estimates cluster in the teens and twenties rather than near the headline lab figure.

The downside is unbounded in a way the upside is not, because one term can invert. Reading an unfamiliar 200-line machine-written diff costs more than reading the 40 lines you would have written yourself, so \(s_{\text{verify}} < 1\) is a normal outcome, not a pathology. With \(f = (0.2, 0.3, 0.35, 0.15)\) and \(s = (1, 3.5, 0.8, 0.9)\), net time is \(0.2 + 0.086 + 0.438 + 0.167 = 0.89\): an 11 percent saving, from a tool that made the authoring phase three and a half times faster. Move \(s_{\text{verify}}\) to 0.65 and the same tool produces a 5 percent loss.

METR's hand-labelled screen recordings show this happening rather than inferring it. With AI available, developers spent less time writing code and less time searching, and more time prompting, waiting for responses, and reviewing output (METR, 2025). The authoring saving was real. It reappeared in three phases the tool does not accelerate.

Why plausible output is the expensive kind

The verification term degrades for a specific reason. Human review heuristics are calibrated on human authorship: attention goes to code that looks rushed, naming that drifts, a comment hedging about an edge case. Generated code strips those cues while keeping a non-zero error rate, so defect locations carry no visual signal and the usual strategy of reading a sample of a large diff becomes uninformative. Practitioners describe this precisely. In Stack Overflow's 2025 survey, the most-cited frustration was output that is "almost right, but not quite", named by roughly two thirds of respondents, with about 45 percent reporting that debugging AI-generated code takes longer; over the same period active distrust of AI output rose to 46 percent from 31 percent (Stack Overflow, 2025, Developer Survey: AI).

The constraint moves downstream

Suppose the authoring improvement survives to the end of the task and more changes get written. They arrive at a review stage whose capacity is set by headcount and attention, and which did not change. Model review as a single-server queue with arrival rate \(\lambda\) and service rate \(\mu\): utilisation is \(\rho = \lambda / \mu\), and expected wait in queue scales as \(\rho / (1 - \rho)\). The non-linearity is the point. A team at \(\rho = 0.70\) absorbing 30 percent more changes moves to \(\rho = 0.91\), where that ratio goes from 2.33 to 10.1, a factor of 4.3 in mean wait, with no change whatsoever in how reviewers work.

Little's law, \(L = \lambda W\), then converts the wait into live branches: more changes, each waiting longer, means substantially more work in progress, which means more conflicts and more rebasing, which consumes the authoring capacity the tool was supposed to free. This is the mechanism behind DORA's instability finding sitting next to its throughput finding.

The counterfactual is eroding

The deepest problem is not any of the above. It is that the comparison itself is becoming unavailable. Measuring an effect requires a world without the treatment, and in a population at 95 percent adoption spending two hours a day with these tools, that world has to be constructed by asking people to work in a way they no longer work. METR's experience is the first clean documentation of what happens next: tasks get withheld from the experiment rather than done without AI, and randomisation into the no-AI arm meets resistance. The resulting selection operates on the task list, which is upstream of everything the design controls.

This is a familiar hazard in a new place. Compliance erosion is routine in clinical trials, and the standard responses (intention-to-treat analysis, non-inferiority designs, stepped-wedge rollouts) exist because blinding fails. Software measurement has not yet imported them. The direction of the resulting bias is not obvious either: the withheld tasks are plausibly the tedious ones where these tools help most, which would push the measured effect downward in whatever work remains.

flowchart TB
  subgraph Favourable["Greenfield task, cheap oracle"]
    A1["New service, no conventions"] --> A2["Author with generation"]
    A2 --> A3["Types and tests decide correctness"]
    A3 --> A4["Large net speedup"]
  end
  subgraph Hostile["Mature repository, human oracle"]
    B1["Legacy module, implicit rules"] --> B2["Author with generation"]
    B2 --> B3["Reviewer reconstructs intent"]
    B3 --> B4["Saving consumed or reversed"]
  end
  A4 --> Agg["Organisation-level average"]
  B4 --> Agg
  Agg --> Useless["Average predicts neither case"]
  classDef emerald fill:#047857,stroke:#34d399,stroke-width:1px,color:#fff
  classDef rose fill:#be123c,stroke:#fb7185,stroke-width:1px,color:#fff
  classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
  class A1,A2,A3,A4 emerald
  class B1,B2,B3,B4 rose
  class Agg,Useless slate

[IMAGE: Two-panel diagram of the same 150-line diff, left annotated as a reviewer sees a human-authored version (inconsistent naming, a hedged comment, an obvious rush), right as a generated version (uniform style, confident naming, one subtle off-by-one). Caption: "The cues reviewers rely on are the cues generation removes."]

Seeing It in Motion

A single change moving through an agent-assisted pipeline shows where each measurement instrument is standing, and why they report different numbers about the same event.

sequenceDiagram
    participant D as Developer
    participant A as Coding agent
    participant CI as CI and tests
    participant R as Human reviewer
    participant P as Production
    D->>A: Task description and context
    A->>A: Generate, run tests, iterate
    A->>D: Candidate change, 240 lines
    Note over D,A: Time-on-task instrument stops here
    D->>CI: Open pull request
    CI->>R: Green build, coverage delta
    Note over R: Throughput instrument counts a merge
    R->>D: Approve after partial read
    D->>P: Deploy
    P-->>D: Incident, 3 days later
    Note over P,D: Delivery instrument finally registers the cost

The three instruments are not measuring the same interval. A trial that stops at the candidate change reports the generator's performance; a telemetry study that counts the merge reports the pipeline's; only a delivery-metrics programme spanning the incident reports the system's, at a lag of days to quarters. That lag is longer than the procurement cycle it is supposed to inform.

stateDiagram-v2
    [*] --> Authored
    Authored --> InReview: pull request opened
    InReview --> Rework: reviewer finds defect
    Rework --> InReview: revised
    InReview --> Merged: approved
    Merged --> Stable: no incident
    Merged --> Incident: defect reaches production
    Incident --> Rework: fix and re-review
    Stable --> [*]
    Incident --> [*]

The loop from Incident back to Rework is the one that generation-era metrics tend to omit, and it is where DORA's instability signal lives. A programme that instruments only the transition into Merged cannot distinguish a team getting faster from a team moving work later.

Watch It Run

Animated diagram of a change flowing from generation through verification into a review queue, with a feedback loop from production incidents back into rework and a self-loop on the agent's generate-and-test cycle.
Solid animated edges carry a single change from task through generation, verification, the review queue and deploy. The animated self-loop on the agent node is its generate-and-test iteration, which is the phase that got cheap. The amber feedback edge from production back into rework is the cost that arrives after every throughput metric has already been recorded. The static Mermaid figures above show the same structure if the animation is absent.

By the Numbers

Study Design Population Outcome measured Effect
Peng et al., 2023 Lab RCT, assigned task Recruited developers, JS HTTP server Time to complete 55.8 percent faster (95 percent CI 21 to 89)
Cui et al., 2025 Three field RCTs, randomised access 4,867 developers at Microsoft, Accenture, a Fortune 100 firm Completed tasks +26.08 percent (SE 10.3); commits +13.55 percent; builds +38.38 percent
Paradis et al., 2024 Enterprise RCT 96 Google engineers, complex task Time on task About 21 percent less time, wide interval
Becker et al., 2025 RCT, self-chosen real tasks 16 experienced maintainers, 246 tasks Time to complete 19 percent slower (CI +2 to +39 percent)
METR, 2026 RCT, redesigned 57 developers, 143 repos, 800+ tasks Time to complete 18 percent faster, CI -38 to +9; new recruits 4 percent, CI -15 to +9
Murphy-Hill et al., 2026 Observational rollout, adopter trajectories Tens of thousands of Microsoft engineers Merged pull requests About +24 percent for adopters, held over four months
Agarwal et al., 2026 Staggered difference-in-differences Open-source repos adopting coding agents Warnings, cognitive complexity About +18 percent warnings, +39 percent complexity (v2; 35 percent in v1)
DORA, 2025 Survey, n in the thousands Industry, 95 percent AI adoption Throughput and instability Both rise with adoption; no offsetting fast repair
Stack Overflow, 2025 Survey 49,000+ respondents Trust and friction 46 percent distrust accuracy, up from 31 percent; about two thirds cite "almost right" output
GitClear, 2025 (via LeadDev) Repository corpus analysis 211 million changed lines, 2020 to 2024 Duplication and refactoring Moved code about 25 percent of lines in 2021 to under 10 percent in 2024; duplicated blocks up eightfold

Sources: effect sizes and intervals are as reported by each study. The Peng et al. sample size commonly quoted as 95 freelance developers appears in secondary coverage and is not confirmed here, so the row omits it. The GitClear figures are from secondary reporting of the company's own report rather than from a peer-reviewed source, and the complexity figure in Agarwal et al. changed between preprint versions, which is noted rather than averaged. Widely circulated vendor telemetry figures for review-time and pull-request-size growth are attributed in different places to different editions of the same report (Faros AI research index); the direction is consistent, the digits are not, so they are excluded from the table.

[IMAGE: Forest plot of the six causal studies, point estimates with confidence intervals on a shared axis of percentage change in time or throughput, zero line marked, each row labelled with design type and population size. Caption: "Six designs, one axis. The intervals overlap zero more often than the headlines suggest."]

A Concrete Example

A platform team of eight engineers adopts an agentic coding tool. Leadership wants a number by the end of the quarter. Here is what each instrument will report, computed from one consistent set of assumptions.

Step 1: baseline phase decomposition. From two weeks of time tracking, an average non-trivial change costs 6.0 hours of engineer time: orientation 1.2 h (20 percent), authoring 1.8 h (30 percent), verification 2.1 h (35 percent), rework 0.9 h (15 percent).

Step 2: apply measured per-phase effects. Generation makes authoring 3.5 times faster, so 1.8 h becomes 0.51 h. Orientation is unchanged at 1.2 h. Verification slows by a factor of 0.8 because diffs are larger and intent is not stated: 2.1 h becomes 2.63 h. Rework slows slightly to 1.0 h.

Step 3: net task time. \(1.2 + 0.51 + 2.63 + 1.0 = 5.34\) hours, against a 6.0 hour baseline. An 11 percent improvement. A time-on-task trial stopping at merge would report roughly that, with an interval wide enough to include zero at this sample size.

Step 4: convert to arrival rate. Eight engineers at 6.0 h per change and 30 productive hours per week produced \(8 \times 30 / 6.0 = 40\) changes per week. At 5.34 h they produce 44.9, call it 45: a 12 percent rise in arrivals, \(\lambda = 9\) per day against the previous 8.

Step 5: the review queue. Reviewers handled 8 per day at \(\mu = 11.4\) per day of effective capacity, so \(\rho\) was 0.70 and the wait factor \(\rho/(1-\rho)\) was 2.33. Larger diffs cut effective capacity to 10.5 per day. Now \(\rho = 9 / 10.5 = 0.857\) and the factor is 5.99. Mean queue wait rises by a factor of 2.6. If a review previously waited 4 hours, it now waits about 10.3.

Step 6: work in progress. By Little's law, open changes rise from \(8 \times 0.5\) days to \(9 \times 1.3\) days, from 4 to about 11.7 concurrent open branches. Three times the live branches across the same eight engineers means merge conflicts and rebasing that did not exist before, charged back to authoring.

Step 7: what each report says. The time-on-task study reports 11 percent faster. The throughput dashboard reports 12 percent more changes merged, and more still if engineers split work into smaller pull requests to clear the queue, which inflates the count without changing delivered behaviour. The delivery metrics report lead time up from 4 h of waiting to 10.3 h, work in progress nearly tripled, and change failure rate up if even one in forty of the thinner reviews lets a defect through.

Every one of those three statements is true. Only the third one is about whether the team is better off, and it is the one that takes a quarter to measure.

[IMAGE: Waterfall chart of the worked example, showing the 6.0 hour baseline decomposed by phase, with the authoring bar shrinking and the verification bar growing, ending at 5.34 hours, with a second panel showing the queue wait rising from 4 to 10.3 hours. Caption: "An 11 percent task-time win and a 2.6 times queue-wait loss, from the same intervention."]

Where It Breaks

Self-report is not weak evidence here; it is evidence of the wrong thing

The usual defence of survey data is that it correlates with the truth. In this domain that has been tested and it failed. METR's participants predicted a 24 percent speedup, measured a 19 percent slowdown, and afterwards still believed they had been sped up by 20 percent, while DORA found more than 80 percent of respondents crediting AI with higher productivity in a population whose delivery instability was rising. The sign is wrong, not the magnitude. A business case resting on "our engineers report saving six hours a week" rests on the one measurement shown not to work.

Throughput metrics are now gameable by construction

Before 2022, counting merged pull requests was a defensible if crude proxy, because writing the code was the expensive part and nobody could inflate the count cheaply. That premise is gone. When generation is near-free, every artefact-count metric becomes an output of the tool rather than an observation of the team, and it improves fastest in the quarter when review queues are deteriorating. This is Goodhart's law with an unusually short latency.

The holdout group has become a retention risk

Running a clean A/B test on tooling requires a group that does not get the tool. With adoption at 95 percent and daily use at around two hours, denial of access is a change to working conditions, and METR's experience shows what follows: refusal, task withholding, and selection on exactly the dimension the design cannot observe. Interrupted time series and stepped-wedge designs are the obvious substitutes, and both are far weaker against confounders than randomisation was.

Heterogeneity makes the organisational average meaningless

The moderators are large and they point in opposite directions. Gains fall with repository familiarity, with codebase maturity, and with verification cost; they rise for less experienced developers and for greenfield work with cheap oracles. The CMU difference-in-differences result adds a sharper version: velocity gains appeared mainly where the agent was the project's first AI tool, with little or short-lived gain where an AI IDE assistant was already in use (Agarwal et al., 2026). An organisation averaging a 40 percent gain on new services against a 10 percent loss on its legacy core learns nothing useful from the mean.

Quality costs are deferred, which means they land on a different decision

Verification can be skipped at merge time; it cannot be skipped forever. The bill arrives as change failure rate, rework, and incidents, which show up one to three quarters after the tooling decision that caused them. Repository-level evidence is consistent with this: static-analysis warnings and cognitive complexity rising after agent adoption, and refactoring signals collapsing across a 211-million-line corpus, with moved code falling from roughly a quarter of changed lines in 2021 to under a tenth in 2024 (reported in DevClass, 2025). None of that is visible in the quarter the licences are bought.

Alternative Designs

Design How it works Key advantage Key limitation Best when
Lab RCT on assigned task Recruit, randomise, time a fixed task Clean causal estimate, fast Task chosen by experimenter; excludes integration Comparing generators on authoring alone
RCT on self-chosen real tasks Randomise tool availability per real task Highest external validity Control arm now suffers refusal and task withholding A population that will still accept randomisation
Field RCT on access Randomise licences inside normal operations Real incentives, large samples Access is not use; effect diluted by non-adopters Estimating organisation-level impact of rollout
Staggered difference-in-differences Compare adopting repos to matched controls over time Thousands of projects, no consent needed Adoption is self-selected; confounded by concurrent change Observing quality trends at ecosystem scale
Interrupted time series with a holdout team One team stays un-tooled, trends compared Survives low willingness to randomise Team-level confounders; holdout becomes a retention risk Organisations that cannot randomise individuals
Delivery metrics programme Track throughput, failure rate, rework, restore time continuously Measures the outcome that matters Correlational; lags the decision by quarters Steady-state operation rather than a procurement decision
Per-change provenance ledger Record AI involvement on every change, then analyse Makes future questions answerable at all Requires discipline before the question is asked Any organisation that intends to measure this seriously

No single design answers the procurement question. The combination that does the most work is a provenance ledger feeding a delivery-metrics programme, with a small assigned-task trial rerun periodically to keep a causal anchor on the authoring phase. That is more instrumentation than most organisations have, and considerably less than the confidence with which they quote percentages.

How It Is Used in Practice

The clearest industrial successes are not general productivity programmes; they are narrow pipelines built where verification is mechanical. Google's migration of 32-bit to 64-bit identifiers across its advertising codebase was driven by an exhaustion deadline rather than an efficiency pitch, and in the reported programme 80 percent of the code modifications in the resulting change lists were purely model-authored, against a target of 50 percent or better acceleration (Nikolov et al., 2025, How is Google using AI for internal code migrations?, ICSE-SEIP 2025, arXiv:2501.06972). A companion industry paper covering 39 migrations run by three developers over twelve months reports 595 submitted changes containing 93,574 edits, of which the model generated 74.45 percent of changes and 69.46 percent of edits, with developers estimating a 50 percent reduction in total migration time (Ziftci et al., 2025, Migrating Code At Scale With LLMs At Google, FSE 2025, arXiv:2504.09691).

The gap between those two percentages is the lesson. Three quarters of the code was generated and half the time was saved, because review and rollout remained human-driven and did not accelerate. Both papers are explicit that the project-specific work lived in the prompts and the validation steps rather than in the model, and that change location was handled by a separate search pass.

The Microsoft study's most practically useful finding is not the 24 percent: first use spread through social networks, and retention tracked engineers' coding activity rather than their demographics. Visible peer use is therefore a rollout lever, and an adoption-based estimate measures the people who chose to keep the tool, not the people who received a licence.

DORA's contribution to practice is its cluster analysis rather than its headline. Only a minority of its seven team profiles were realising throughput gains without a rise in change failure rate, which locates the difference in surrounding practice: small changes, fast reliable tests, loose coupling, and review capacity planned as deliberately as compute.

[IMAGE: Pipeline schematic of the Google migration architecture: change-location pass feeding per-site prompts, a validation stage running build and tests with a retry loop, then batching into reviewable change lists routed to owners, with the human stages shaded to show where the time savings stop. Caption: "Generation is three quarters of the edits and half of the time saved."]

Insights Worth Remembering

  1. The spread in the evidence is a feature of the outcome variables, not a failure of the studies. Time-on-task, artefact throughput and delivery outcomes answer different questions, and AI moves them in different directions at once, so ask which one a quoted percentage measured before arguing about its size.

  2. Generation-only improvements have a low ceiling and no floor. Authoring is typically 25 to 40 percent of a change, so an infinitely fast generator saves at most that fraction, while a degraded verification phase can push net time above baseline. The verification term sets the sign.

  3. Mechanical oracles are the enabling technology, not the model. Strict types, fast deterministic tests and differential comparison are what convert cheap generation into shipped change, which is why the best industrial results come from tasks where a compiler decides correctness.

  4. Capacity moves, it does not vanish. A 30 percent arrival increase against a review stage at 70 percent utilisation roughly quadruples mean queue wait, and teams that scale review usually find the constraint relocate to integration or release coordination rather than disappear.

  5. Self-report has been falsified in this specific domain. Participants who were measurably 19 percent slower believed they were 20 percent faster, so treat developer sentiment as a measure of adoption and morale rather than of output.

  6. Artefact counts became endogenous in 2022 and nobody retired them. Pull requests, commits and lines are now partly outputs of the tool under evaluation, so any metric programme lacking a paired stability measure will report success precisely when it should report risk.

  7. The counterfactual is a depreciating asset. Each year of deeper adoption makes the no-AI arm harder to populate and less representative, so take a causal estimate now and record per-change provenance regardless: it is the only instrument whose value grows with time.

Open Questions

Does the heterogeneity converge as agents get better, or is it structural? Measured: gains fall with repository familiarity and maturity across multiple designs, and the CMU study found gains concentrated on a project's first AI tool. Unknown: whether that reflects a capability gap that closes, or a permanent property of work whose cost is dominated by human verification. The two hypotheses make opposite predictions about 2027 and are distinguishable by rerunning a fixed task battery across model generations.

What is the right denominator for complexity growth? Rising cognitive complexity and duplication are measured; complexity per unit of delivered behaviour is not, because nobody can quantify the denominator. Until that changes, maintainability claims in either direction rest on proxies, and more repository mining will not settle them.

Can stepped-wedge or non-inferiority designs replace randomisation here? They are the standard responses to compliance erosion in clinical research, and nothing in the software literature has imported them yet. Whether they survive the temporal confounding of monthly model releases is open.

How much of the measured instability is deferred verification rather than worse code? DORA's correlation is between adoption and instability; the mechanism could be defect density in generated code, the same code reviewed less carefully, or simply more changes per unit time. These have different fixes, and separating them needs per-change provenance that almost nobody records.

Does the throughput gain survive once review is properly resourced? It is plausible that teams which expand review capacity in proportion to generated volume capture most of the authoring saving as delivered change. That is an inference from queueing theory, not a measured result, and it is probably the single highest-value experiment available to a large engineering organisation right now.

Sources and Further Reading

  1. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." METR. arXiv:2507.09089; summary at metr.org
  2. METR (2026). "We are Changing our Developer Productivity Experiment Design." metr.org/blog/2026-02-24-uplift-update
  3. Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot." arXiv:2302.06590
  4. Cui, Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers." Management Science. SSRN 4945566
  5. Paradis, E., et al. (2024). "How much does AI impact development speed? An enterprise-based randomized controlled trial." ICSE-SEIP 2025. arXiv:2410.12944
  6. Murphy-Hill, E., Butler, J., & Savelieva, A. (2026). "Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI." arXiv:2607.01418
  7. Agarwal, S., He, H., & Vasilescu, B. (2026). "AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development." MSR 2026. arXiv:2601.13597
  8. Nikolov, S., Codecasa, D., Sjövall, A., Tabachnyk, M., Chandra, S., Taneja, S., & Ziftci, C. (2025). "How is Google using AI for internal code migrations?" ICSE-SEIP 2025. arXiv:2501.06972
  9. Ziftci, C., Nikolov, S., Sjövall, A., Kim, B., Codecasa, D., & Kim, M. (2025). "Migrating Code At Scale With LLMs At Google." FSE 2025 Industry. arXiv:2504.09691
  10. DORA / Google Cloud (2025). "State of AI-Assisted Software Development." PDF; overview at blog.google
  11. Stack Overflow (2025). "Developer Survey 2025: AI." survey.stackoverflow.co/2025/ai
  12. GitClear (2025). "AI Copilot Code Quality" research, covered in LeadDev and DevClass
  13. Faros AI. Research index of engineering telemetry reports. faros.ai/research (vendor research; figures differ between editions)
  14. Library concepts: measuring AI-assisted developer productivity, the verification cost of generated code, reviewing machine-authored changes, LLM-assisted code migration at scale, AI-assisted code and maintainability signals

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.