Stacked Changes
also called Stacked Diffs, Patch Series, Dependent Change Chain
Splitting one body of work into a chain of small dependent changes that each review and merge on their own, so branch lifetime is set by the size of a link rather than the size of the feature.
A team moves to trunk-based development and branch lifetime does not fall. The reason is arithmetic: elapsed time from first commit to merge is time writing + time waiting for a first review + time in review rounds + time in the merge queue, and the branching model only touches the last term.
A 1,400-line change waits 9 hours for a first look and takes three rounds. Nothing about calling the branch short-lived makes a reviewer read it faster. The constraint was the review shape and the team changed the release shape. Stacked changes attack the review shape directly: the feature becomes five dependent changes of a few hundred lines each, independently reviewable, each merged on its own behind a flag.
Why it matters
Review effectiveness falls off sharply with size. Common review guidance converges on roughly 200 to 400 lines as the point beyond which reviewers stop finding defects and start approving, and that number has nothing to do with the branching policy. A stack keeps every reviewed unit under that ceiling without forcing the feature to be smaller.
The second effect is trunk health: a five-day branch is five days of unintegrated divergence, and the longer it lives the more likely its merge breaks something. The third decides adoption, because a reviewer asked for 300 lines responds the same day and the same reviewer asked for 1,400 lines schedules it.
Implementation patterns
- One logical change per link, ordered so each is correct alone: a refactor, the new interface, the implementation, the call-site switch, then deletion of the old path, with a flag at the top of the stack so early links do not expose a half-built feature.
- Automated restack. When a lower link merges, every link above it must be rebased. A five-deep stack with one merge a day is four rebases a day, roughly 45 minutes of manual work plus conflict mistakes, which is why manual stacking dies inside a month.
- Per-link diffs in review tooling, or reviewers read the same code five times.
- Merge queue per link, so each is tested against the trunk state that will exist after it merges.
Industry example
The Linux kernel has used this model since the 1990s, long before tooling existed for it. Work is posted as a patch series: each patch is self-contained with its own message, reviewed individually, and required to leave the tree building and bootable so bisection stays meaningful. Maintainers routinely accept the first six patches of a ten-patch series and send the rest back.
Gerrit encoded the same model in a tool, where one commit is one change and changes form dependent chains each carrying its own review state. Both show the real prerequisite, which is not discipline but tooling that makes restacking free. Bisectability is the part teams skip and the part that pays off in production, because a link that does not build alone makes the history useless.
Failure scenarios
- Links that are not independently correct, so a bisect lands on a commit that does not build.
- A stack ten links deep, where lower links sit in review long enough that upper ones are rewritten against a moving base.
- No flag, so link three exposes a feature with no UI and support fields the tickets.
Trade-offs
Stacking buys short-lived links and effective review, and pays in tooling and author overhead. The author does more work splitting and maintaining the chain so the reviewer does less, which is the right trade when reviewer time is the constraint and the wrong one when it is not. It also raises the merge count, so queue throughput matters more: five links at a 20-minute pipeline is 100 minutes of queue per feature, and a serial queue can become the new ceiling.
When not to use it
A change that genuinely fits in 300 lines needs one pull request, not a stack. Stacking a small change adds coordination for nothing.
Skip it where reviewer wait is not the constraint: a two-person team reviewing within the hour gains nothing. And with no automated restack, choose smaller independent pull requests instead, because a badly supported stack is slower than one large review.
Interview question
Q: "Your organisation mandates trunk-based development and the mean time from first commit to merge is still four days. Walk me through how you would find out why, and what you would change."
What a strong answer covers: decomposing elapsed time into the four terms and showing the model only touched the queue; measuring time to first review and change size before proposing anything; identifying a small approver group as a service-rate constraint; proposing stacked changes with restack tooling as a prerequisite rather than an optimisation; the merge-queue throughput consequence; and the limit, that if branches are long because features ship unflagged, the flag comes first.
Quick check
Quiz: Why does a stack of five 200-line changes shorten branch lifetime when a single 1,000-line change does not? — Because lifetime is set by the review round trip on the reviewed unit, and a unit a reviewer can finish in one sitting comes back in hours rather than days.
Flashcard: What decides whether stacked changes survive? — Automated restacking. Four manual rebases a day is roughly 45 minutes plus conflict risk, and teams abandon the practice within a month without it.