beginner 2 min answer Multiple choice

A team reviews small changes carefully and approves large ones within minutes. Why does defect detection collapse as a change grows, and which single intervention recovers the most of it?

code reviewdefect detectionbatch sizeattentiondeveloper experience
Pick one
Show the full answer Hide the answer

The mechanism

Review finds defects by reading code while holding the surrounding context in your head, and the attention available for one review session is roughly fixed. As the diff grows, the reader's model of the change degrades, the reading rate rises to cover the ground, and past some point reading becomes skimming. The approval still happens; the inspection does not.

The most-cited published study of peer review — about 2,500 reviews covering roughly 3.2 million lines of code, published in 2006 — put numbers on it: defect detection is strongest at roughly 200 to 400 changed lines per sitting, falls sharply above a reading rate of about 500 lines an hour, and degrades after 60 to 90 minutes of continuous review. Treat those as orders of magnitude rather than targets, and the shape still holds in any codebase.

The second mechanism, which is social

A large change is also harder to object to. By the time 2,000 lines exist, asking for a different approach costs the author days of visible rework, and reviewers know it. So reviews of large changes drift toward comments on naming and away from comments on design, which is the opposite of what review is for. Splitting the work restores the reviewer's ability to say no while saying no is still cheap.

Why the other options fail

  • A third approver. Three skimmers find less than one reader. Adding reviewers to an oversized change also diffuses responsibility, because each assumes someone else looked closely, and it adds a third waiting queue to lead time. It is a reasonable control for a narrow class of high-risk change, not a fix for comprehension.
  • A checklist. Checklists work for known, enumerable classes: migrations, secrets, authorisation checks. They do nothing for understanding, and on a 2,000-line diff the boxes get ticked. Keep the checklist and stop expecting it to carry the review.
  • A coverage threshold. Coverage records that a line executed, not that anything asserted on its behaviour. Raising the bar on large changes mostly produces tests written against the implementation, which then block the refactor the next reader wants.

When not to split a change

Mechanical diffs. A generated or rewritten change across 1,100 files should be reviewed by reading the generator and confirming the output is reproducible, not by reading the output. Lumping those with hand-written changes is how teams conclude that the size rule is unrealistic.

The other case is a change that arrives as a readable sequence of commits, each one a complete thought. That is the same intervention as splitting, performed by the author instead of enforced by the process.