practice

Merge Queue

also called Merge Train, Serialised Integration

Testing each change against the combination that will actually exist after merge, rather than against the base it was written on - preventing the broken trunk that concurrent merges produce.

continuous-integrationtrunktestingthroughputsemantic-conflicts

Two changes each pass their tests against the current trunk, and both are merged. Neither was tested against the other, and together they break the build — a semantic conflict, where each change is individually correct and the combination is not.

At low merge rates this is rare. At high rates it is routine, and it produces a trunk that is broken for reasons no individual author could have prevented.

A merge queue serialises the final verification: each change is tested against the current trunk plus every change ahead of it in the queue, and merges only if that combination passes.

Why it matters more as merge rate rises

The probability of a semantic conflict grows with the number of changes merged in the window between a change being tested and being merged. Above a certain rate, a shared trunk without a queue is broken more often than it is working — which destroys the value of continuous integration entirely, because feedback stops being trustworthy.

Implementation patterns

  • Speculative parallel execution. Testing each queue position sequentially serialises throughput to one change per pipeline duration. Testing several prospective combinations in parallel, and discarding the ones invalidated by a failure ahead, recovers most of the throughput.
  • Batching with bisection. Test several changes together; on failure, bisect to find the offender rather than rejecting the batch. Much higher throughput at the cost of a slower failure path.
  • Automatic ejection of a failing change with a clear message, so the queue continues rather than stalling on one bad change.
  • Priority classes, so an urgent fix is not queued behind a routine change.
  • A queue length and wait-time metric, since the queue becomes the delivery bottleneck if the pipeline is slow — the queue makes pipeline speed matter more, not less.

Industry example

Large monorepos with many contributors adopt this because the alternative is a permanently red trunk. It pairs naturally with change-aware test selection: the queue verifies the combination, and selection keeps each verification affordable.

The interaction is worth noting. Test selection makes individual verification fast; the merge queue makes the result trustworthy. Selection without a queue risks merging combinations nobody tested; a queue without selection is too slow to keep up with the merge rate.

Failure scenarios

  • Sequential queue execution with a slow pipeline, making the queue the constraint and lengthening lead time for everyone.
  • No automatic ejection, so one failing change stalls the queue.
  • Batching without bisection, rejecting good changes alongside the bad one and eroding trust.
  • Flaky tests in the queue, which reject valid changes and train people to force-merge — at which point the queue provides nothing.
  • Bypass paths used routinely, which reintroduces the problem the queue exists to solve.

Trade-offs

A merge queue adds latency between approval and merge, which developers feel. It also requires pipeline capacity for speculative execution, and it makes pipeline speed and flakiness far more consequential.

For a repository with few merges per day it is unnecessary complexity — testing against the base is sufficient when the window is small. It earns its cost when the merge rate makes semantic conflicts routine.

Interview question

"Two changes each pass CI and together break the trunk. Explain how, tell me how you prevent it, and tell me what that prevention costs at fifty merges an hour."