A repository takes 420 merge attempts a day. The merge queue validates one change at a time against the current tip and the suite takes 22 minutes. Developers now wait hours to land anything. Which change raises landed merges per day the most?
Show the full answer Hide the answer
The deciding property
A strictly serial queue has a throughput ceiling that no amount of hardware touches. One validation at a time at 22 minutes each gives 1,440 / 22, about 65 merges a day, against 420 demanded. The gap is nearly sevenfold, so only a change that attacks the serialisation itself can close it.
Batching does exactly that. Validate the next n queued changes as one combined candidate; if it passes, all n land from one pipeline run. The speculative form goes further and starts the n+1 candidate optimistically on the assumption that n passes, so the pipeline is never idle waiting for a verdict.
The arithmetic that sets the batch size
Batching trades throughput against the cost of a failure. If each change independently breaks the combined build with probability p, a batch of n fails with probability 1 minus (1 minus p) to the power n.
- At p = 4%, a batch of 8 fails about 28% of the time.
- At p = 4%, a batch of 20 fails about 56% of the time, and every failure costs a bisection of roughly log2(n) extra runs.
The decision rule: raise n until the batch failure rate approaches roughly 30%, then stop raising n and start lowering p. Beyond that point the bisections eat the gain, and the lever that matters is the rate at which changes arrive broken, which is mostly flaky tests. This is the useful inversion: at high batch sizes, flake rate is a throughput constraint rather than a credibility problem.
Why the other options fail
- Triple the runners. It does nothing while validation is serial, because the constraint is the dependency between attempts, not the supply of machines. It becomes the right answer the moment batching or speculation is in place, which is why these two are usually bought together.
- Only the affected tests per attempt. Genuinely cuts the suite and genuinely raises the ceiling, and it is unsafe exactly here. A merge queue exists to catch the interaction between concurrently queued changes, and per-change test selection is computed from one change's dependency graph, so it cannot see that interaction. Use selection on pre-merge pull requests and the full suite in the queue.
- Skip revalidation for young branches. This deletes the queue's only purpose. A branch a day old was tested against a tip that 65 merges have since replaced, and the broken trunk it causes costs every engineer in the repository more than the wait it saved.
When a merge queue is the wrong tool
Below roughly 20 merges a day, the queue adds latency to every change to prevent a broken trunk that happens twice a month. Post-merge CI with a fast revert and an owner who is expected to use it is cheaper and teaches the same discipline. The queue earns its latency when the trunk breaks often enough that the cost of a break, multiplied by everyone blocked, exceeds the cost of making every change wait.