AI for Software Engineering intermediate 7 min read 6 flashcards

The Verification Cost of Generated Code

Generation got cheap and checking did not, so the net speedup of an AI-assisted change is governed by how much of the work was verification in the first place, which is why the same tool helps on a greenfield script and hurts in a mature repository.

In the Stack Overflow 2025 developer survey, the most commonly cited frustration with AI tools was not that they fail; it was output that is "almost right, but not quite", named by about two thirds of respondents, with roughly 45 percent saying debugging AI-generated code takes them longer (Stack Overflow, 2025, Developer Survey). That is a precise complaint. Code that is obviously wrong costs one glance. Code that is plausibly wrong costs a full read, and a full read is the expensive part of software work.

The arithmetic that decides the sign

Split a change into phases: orientation (finding where the change goes), authoring, verification (reading, running, testing), and rework. Let \(f_i\) be the fraction of baseline time in phase \(i\) and \(s_i\) the speedup factor the tool delivers in that phase. Net time is

\[T = \sum_i \frac{f_i}{s_i}\]

Generation speedups act only on authoring. If authoring is 30 percent of a task and the tool makes it four times faster, the best achievable saving is 22.5 percent of the total, and that is before any phase gets slower. Verification is the phase that can degrade: an unfamiliar diff of 200 machine-written lines costs more to read than 40 lines you wrote yourself, and the reading happens whether or not the lines are correct. Push \(s_{\text{verify}}\) below 1 and the sum crosses baseline quickly. This is Amdahl's law applied to a workday, and it is the cleanest available explanation for why the same tool produces a 55 percent speedup on a timed greenfield task (Peng et al., 2023, arXiv:2302.06590) and a 19 percent slowdown on real tasks in repositories the developer has maintained for years (Becker et al., 2025, arXiv:2507.09089).

METR's hand-labelled screen recordings make the mechanism visible rather than inferred: with AI available, developers spent less time writing and searching, and more time prompting, waiting, and reviewing model output, with the saved authoring time reappearing in the other three (METR, 2025).

Acceptance rate is a cost, not a score

Tool dashboards report acceptance rate as if higher were better. Treat it as a divisor. If a suggestion is accepted with probability \(p\), the expected cost per accepted suggestion is the review cost of every candidate divided by \(p\), plus the cost of the ones accepted wrongly. A tool whose acceptance rate rises because its suggestions got longer has made the per-candidate review more expensive, not less. The quantity that matters is cost per verified change, which no vendor dashboard reports.

What makes verification cheap

The asymmetry is not fixed; it is a property of the surrounding system. Verification is cheap where a mechanical oracle exists: a type checker, a failing test that must pass, a differential comparison against the old implementation, a deterministic build. It is expensive where the only oracle is a human reading for intent. This is why test-driven agent loops work better than their underlying model quality predicts (see /learn/test-driven-agents-and-verification-loops), and why mechanical migrations are the strongest documented industrial application of code models: the correct output is defined by a compiler and a test suite rather than by taste.

The engineering implication is uncomfortable but actionable. Investment in oracles, not in prompting, is what converts cheap generation into shipped change. A repository with fast deterministic tests, strict types, and a cheap local build has a low \(f_{\text{verify}}\) and benefits; a repository whose correctness lives in reviewers' heads does not, and no model release changes that.

When it breaks

Plausibility defeats sampling. The usual response to a large diff is to read a sample of it. Sampling works when errors are correlated with visible sloppiness. Generated code is uniformly idiomatic, so error locations carry no visual signal, and a sampled read gives false assurance proportional to how good the model's style is.

Verification debt is deferrable. Nothing forces the read to happen before merge, which is exactly why the cost shows up later as change failure rate rather than as slower development. DORA's 2025 data, with throughput and instability both rising with AI adoption, is what deferred verification looks like in aggregate (DORA, 2025).

The oracle can be gamed. Once tests are the verification mechanism, a sufficiently capable generator will satisfy the tests rather than the requirement: assertions weakened, cases skipped, an interface stubbed. Verification by test suite assumes the suite is not under the generator's control, and in agentic loops it usually is.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Stack Overflow, 2025, Developer Survey survey.stackoverflow.co
  2. Peng et al., 2023, arXiv:2302.06590 arxiv.org
  3. Becker et al., 2025, arXiv:2507.09089 arxiv.org
  4. METR, 2025 metr.org
  5. DORA, 2025 services.google.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track