AI-Assisted Code and Maintainability Signals
The measured side effects of assistant-heavy development, covering duplication, the collapse of refactoring signals, rising complexity and delivery instability, and which of these signals are strong enough to act on.
In GitClear's analysis of 211 million changed lines from 2020 to 2024, code that was moved rather than added fell from roughly a quarter of changed lines in 2021 to under a tenth in 2024, and 2024 became the first year in their dataset in which copy-pasted lines exceeded moved lines, alongside an eightfold rise in blocks of five or more duplicated lines (reported in DevClass, 2025 and LeadDev, 2025). Moved code is the fingerprint of refactoring. Its disappearance is the fingerprint of a workflow in which adding a new block is cheaper than finding the existing one.
Why the mechanism is structural
An assistant optimises the next edit, not the repository. Three properties of the generation loop push toward duplication rather than reuse. The model's view of the codebase is a retrieved window, so a function that already does the job is frequently not in context (see /learn/repository-context-and-the-retrieval-problem). Writing a fresh implementation is a single accepted suggestion, whereas reusing one requires locating it, reading its contract, and possibly changing it, which is work the loop does not reward. And the accepted-suggestion feedback signal is local: nothing in the interaction observes that the repository now contains the fourth copy of a date parser.
The quasi-experimental evidence points the same way. A difference-in-differences study of coding-agent adoption across open-source repositories found static-analysis warnings up by roughly 18 percent and cognitive complexity up by roughly 39 percent after adoption, with the complexity figure revised between preprint versions (35 percent in v1), while velocity gains were confined largely to projects for which the agent was their first AI tool (Agarwal, He and Vasilescu, 2026, AI IDEs or Autonomous Agents?, MSR 2026, arXiv:2601.13597).
The delivery-level signal
Repository metrics describe the artefact; delivery metrics describe the consequence. DORA's 2025 survey found AI adoption correlating with higher delivery throughput and, at the same time, with higher instability: more change failures, more rework, longer recovery. The report explicitly tested the optimistic hypothesis, that teams shipping faster also repair faster and therefore come out even, and found no support for it (DORA, 2025). Its cluster analysis matters more than the headline: only a minority of team profiles were realising throughput gains without a rise in change failure rate, which makes the practice, not the tool, the variable.
Which signals to trust
Not all of these measures are equally sound, and treating them as equally sound is how a metrics programme loses credibility.
Reasonably trustworthy: duplication of non-trivial blocks (mechanically detectable, low false-positive rate), change failure rate and rework, cognitive complexity per changed function, and time from open to merge as a congestion signal. Weak: lines of code in either direction, commit counts, acceptance rate, and any vendor-supplied composite "productivity index". Vendor telemetry reports in this area also disagree with themselves; widely quoted figures for a 91 percent increase in review time and a 154 percent rise in pull-request size are attributed in different places to a 2025 and a 2026 edition of the same report, which is reason enough to cite the direction and not the digits (Faros AI, research index).
A defensible programme instruments three things together: a throughput measure, a stability measure, and a maintainability measure, reported on the same cadence, with AI involvement recorded per change so the comparison is possible at all. Provenance is the part teams skip and then cannot reconstruct.
When it breaks
Duplication is not always worse. A deliberately copied adapter at a module boundary can be better than a shared abstraction that couples two services. The signal to act on is duplication that crosses a correctness boundary, where one copy will be fixed and the others will not.
Attribution is unavailable after the fact. Once assistant-written and hand-written code are interleaved in the same commits, no later analysis can separate them cleanly, and repository-level before-and-after comparisons absorb every other change the team made in that period, including headcount and product mix.
Complexity growth can be the price of delivered scope. A codebase that grew more complex while shipping twice the functionality has not necessarily degraded. The honest form of the claim is complexity per unit of delivered behaviour, which is exactly the quantity nobody can measure, which is why this literature leans on proxies and should be read with that in mind.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- DevClass, 2025 devclass.com
- LeadDev, 2025 leaddev.com
- Agarwal, He and Vasilescu, 2026, AI IDEs or Autonomous Agents?, MSR 2026, arXiv:2601.13597 arxiv.org
- DORA, 2025 services.google.com
- Faros AI, research index faros.ai
6 flashcards for this concept
Click a card to reveal the answer.