AI for Software Engineering advanced 7 min read 6 flashcards

Reviewing Machine-Authored Changes

When generation gets cheap, human review becomes the binding constraint on delivery, and the queue behaves the way queues do: larger diffs, longer waits, and a reviewer whose attention is the scarce resource.

A team that merges 24 percent more pull requests has not necessarily shipped 24 percent more value, but it has certainly created 24 percent more review (Murphy-Hill et al., 2026, arXiv:2607.01418). Review capacity is fixed by headcount and attention, it did not change when generation got cheap, and it is now the stage where most of the queueing happens.

The queue, with numbers

Treat review as a single-server queue. With arrival rate \(\lambda\) changes per day and service rate \(\mu\) reviews per day, utilisation is \(\rho = \lambda/\mu\) and expected waiting time grows as \(\rho/(1-\rho)\). The non-linearity is the whole point. A team at \(\rho = 0.7\) that increases authoring throughput by 30 percent moves to \(\rho = 0.91\), and mean wait in queue rises roughly fivefold. Nothing about the reviewers changed; the system just moved up the curve.

Two second-order effects make it worse. Service time per change rises when diffs get larger, which is the direction generated changes move. And the reviewer's effective service rate falls with fatigue, which is not a metaphor: a reviewer on their ninth idiomatic-looking diff of the day is not applying the attention they applied to the first.

The practical consequences follow from Little's law, \(L = \lambda W\). Work in progress grows with both arrival rate and wait, which means more branches live longer, which means more merge conflicts and more rebasing, which consumes authoring capacity that the assistant was supposed to have freed.

Why machine-authored diffs are harder to review, not easier

The usual review heuristics are calibrated on human authorship. Reviewers spend attention where the code looks rushed, where the author flagged uncertainty, and where naming is inconsistent. Generated code removes all three cues while leaving the error rate non-zero; it is uniformly well-formatted, confident, and conventionally named. Stack Overflow's 2025 survey puts the practitioner version of this plainly: output that is almost right but not quite was the most-cited frustration, named by about two thirds of respondents (Stack Overflow, 2025).

There is also an authorship asymmetry. A human author can answer "why did you do it this way?"; that answer is the cheapest form of review there is. When nobody can answer it, the reviewer must reconstruct the intent from the diff, which is strictly more work than confirming a stated intent.

Automation bias compounds both effects. Reviewers accept recommendations from a system they believe to be competent at a higher rate than the system's accuracy justifies, and explanations attached to a recommendation raise acceptance of wrong recommendations as well as right ones (see /learn/appropriate-reliance-and-its-two-failures).

What actually relieves the constraint

Adding reviewers is the expensive answer and usually unavailable. The cheaper interventions act on the three terms in the queue.

Reduce arrival size: enforce small, single-purpose changes, which is harder to hold when the marginal cost of writing more is near zero. Raise the service rate with mechanical pre-review: types, linters, coverage deltas, static analysis and model-based review as a filter that removes classes of defect before a human looks, not as a substitute for the human (see /learn/code-review-by-model). Reduce arrivals that should not have been authored at all, which is a planning problem rather than a tooling one. And record provenance on each change, because a reviewer who knows which hunks were generated can spend attention accordingly.

When it breaks

Model-based review raises throughput and lowers the floor. A filter that catches the easy defects also trains reviewers to expect clean diffs, which is precisely the expectation that makes the hard defects pass. The failure is silent and shows up as change failure rate.

Rubber-stamping is a rational response to a full queue. When the backlog is visibly unmanageable, approval becomes the cheapest way to clear it. Any review programme that does not measure read time per change cannot distinguish this from working review.

The bottleneck relocates rather than clearing. Teams that successfully scale review frequently find the constraint moves to integration testing, staging capacity, or release coordination. Whether total delivery improved is then a question about the slowest stage, not the fastest, which is why DORA's instability finding sits alongside its throughput finding rather than contradicting it (DORA, 2025).

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Murphy-Hill et al., 2026, arXiv:2607.01418 arxiv.org
  2. Stack Overflow, 2025 survey.stackoverflow.co
  3. DORA, 2025 services.google.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track