Keeping the main branch green: thirteen years of merge queues, read from the repositories that ran them
How large projects stopped trusting 'tested at review time' and built merge queues: the not-rocket-science invariant, speculative chains versus human-curated rollups, and when to take the platform's queue versus owning the bot.
The merge-queue pattern traced through its primary record: Rust's three generations of bors, Zuul's speculative gating, Kubernetes' Tide, GitLab's trains, Chromium's CQ and GitHub's built-in queue, plus two committed postmortems and git-history measurements of what the queues actually land. A reader leaves able to size a queue against CI duration and merge rate with a worked formula, name the four production failure classes, and argue the platform-versus-bespoke decision with the recorded reasons on both sides.
After GitHub absorbed the merge queue into the platform and Bors-NG deprecated itself in response, the project that invented the pattern rebuilt its own bot again, and the 2026 rewrite still tests exactly one candidate at a time: measured September 2026 history shows its throughput comes from humans packing 92% of PRs into risk-graded rollups, not from machine speculation.
What you get out of it
- A serial queue's viability is (arrivals x CI-hours) / (PRs-per-build x 24) < 1; rust-lang/rust closes a 4-5x gap purely with human-curated rollups averaging 9.7 PRs per build (measured from git, Sept 2026).
- Speculative parallel queues (Zuul, GitLab, GitHub) all degrade to serial testing plus discarded builds when failures cluster; Zuul's own docs state the worst case, and GitLab documents the restart cascade.
- Third-party queues die on platform integration seams, not algorithms: Bors-NG's deprecation notice enumerates bugs only the platform could fix, while Rust's satellite repos lost reviewer delegation by adopting the platform queue and rebuilt it in a bot.
- The queue multiplies CI availability into merge availability: a kernel.org mirror outage closed the rust-lang/rust tree for ~4 days, and the postmortem's model shows five external dependencies at 99% uptime push CI below 95%.
- No public postmortem in this corpus records the invariant itself failing (a tested merge turning main red); every recorded failure is an availability failure, including the merge bot taking itself down for ~8 hours with a bug both its test layers missed.
Scope
Why this, now. Rust completed its migration off eleven-year-old Homu in Q4 2025 and spent 2026 operating the rewrite, so for the first time the full arc, invention (2013), ecosystem adoption, platform absorption (2023) and the inventor's decision to rebuild anyway, is readable end to end in public repositories.
What it does not cover. Monorepo-internal systems with no public repository trail (Google TAP, Meta), commercial queue products, CI cost optimisation as its own topic, and anything hosted outside GitHub, because this session's network policy reached GitHub only; the EuroSys 2019 SubmitQueue paper and all conference talks are acknowledged but uncited for that reason.
Other field guides
Scaling without splitting: a decade of Shopify, read from the code it published
A decade of one company's architecture, reconstructed from the ring of repositories around a closed monolith: a boundary checker whose stricter half …
26 sources · 8 organisations · 4 postmortemsTwo hundred control planes per cluster: ten years of SAP's Gardener
A fleet platform has to give hundreds of teams their own cluster, on whichever infrastructure each product sells on, cheaply enough that asking for o…
34 sources · 4 organisations · 3 postmortemsThe parts that outlived the product: ten years of Docker, read from its own repositories
A decade of one company's architecture told through the sequence every platform team eventually faces: bundle to win the workflow, extract components…
38 sources · 9 organisations · 5 postmortems