Keeping main green  / field guide
Practitioner field guide · 2026-10-04

Keeping the main branch green

Thirteen years of merge queues, reconstructed from the repositories that ran them: Rust's three generations of bors, OpenStack's Zuul, Kubernetes' Tide, GitLab's trains, Chromium's commit queue and GitHub's built-in feature. After reading it you can size a queue against your own CI duration and merge rate, name the four ways queues fail in production, and say when to take the platform's queue and when to own one.

30 ledger entries, every one fetched 9 organisations 3 postmortems Evidence through Oct 2026 Read: ~25 min
01

The territory

A write-serialisation problem wearing a developer-workflow costume: the shared branch is a contended resource, and testing before review stops protecting it the moment test results go stale faster than they are produced.

State the problem without naming a tool. Many writers submit changes to one shared line of history. Each change is verified against a snapshot of that history, but other changes land between verification and commit, so the state that was verified is not the state that ships. Two changes can each pass alone and break together; the public record's canonical example is one pull request renaming bifurcate() while another adds a call to the old name, both green in isolation, red the moment both land [5]. When that happens the breakage lands on everyone at once: the shared branch is the one dependency every developer pulls.

The fix has been restated by every generation of tooling since Graydon Hoare wrote it down for the early Rust project: automatically maintain a repository of code that always passes all the tests [24]. Mechanically that means the thing you test is the exact merge result you are about to publish, and nothing reaches the shared branch any other way. Zuul's documentation says it in one sentence: a gating system "should always test each change applied to the tip of the branch exactly as it is going to be merged" [1]. Everything else in this guide, batching, speculation, rollups, bisection, is throughput engineering layered on that one invariant, because the naive implementation serialises the whole organisation behind one CI run.

737
PRs landed on rust-lang/rust in Sept 2026, all through one serial test lane (measured from git)
~3.5 h
One full queue build of rust-lang/rust, per the project's own rollup procedure
92%
Share of those September PRs that rode a human-assembled rollup rather than landing alone (measured)
4 days
rust-lang/rust tree closed in July 2026 when a kernel.org mirror outage turned the queue's CI red

Who has solved this in production, on the public record: the Rust project across three generations of its own bot (bors 2013, Homu 2014, a Rust rewrite merging since late 2025) [2][3][16], OpenStack with Zuul's dependent pipelines [1], Kubernetes with Tide since 2017 [19], the hosted Bors-NG instance that served the wider GitHub ecosystem from 2017 to its 2023 deprecation [5][9], Chromium's commit queue [25], GitLab's merge trains [23], Smarkets' marge-bot for GitLab [24], and GitHub itself, whose built-in merge queue absorbed the pattern into the platform [22].

The surprise, up front. The story was supposed to end in 2023. GitHub shipped a built-in merge queue, and Bors-NG, the tool that popularised the pattern, deprecated itself the same spring, its maintainer writing that the platform can fix "a bunch of bugs in bors-ng that we can't fix" [7]. Instead, the project that invented the pattern went the other way: between 2025 and 2026 Rust completed an in-house rewrite of its queue [16][17], and the rewrite still tests exactly one merge candidate at a time, by design [14]. Its throughput lever is not machine speculation but a social mechanism: humans risk-grade every PR and pack the low-risk ones into rollups, and the September 2026 numbers say 92% of landed PRs travelled that way. The most sophisticated user of merge queues on GitHub treats the queue as a people system with a bot in the middle, not a scheduling problem.

Figure 1 · Thirteen years of the same invariant

same invariant

2013 · graydon/bors
one PR at a time, cron loop
(mozilla/rust, Buildbot)

2014 · Homu
state + webhooks, try builds,
priorities, rollup command

2017 · Bors-NG
batching with bisection,
hosted for the GitHub ecosystem

2017 · Tide
replaces k8s Submit Queue;
search-driven, batches

≤2019 · Zuul v3 (OpenStack)
speculative dependent pipelines,
the parallel branch of the family

2023 · GitHub merge queue ships;
Bors-NG deprecates itself

2025–26 · Rust rewrites bors;
satellite repos take the platform queue,
rust-lang/rust keeps its own serial lane

same invariant

2013 · graydon/bors
one PR at a time, cron loop
(mozilla/rust, Buildbot)

2014 · Homu
state + webhooks, try builds,
priorities, rollup command

2017 · Bors-NG
batching with bisection,
hosted for the GitHub ecosystem

2017 · Tide
replaces k8s Submit Queue;
search-driven, batches

≤2019 · Zuul v3 (OpenStack)
speculative dependent pipelines,
the parallel branch of the family

2023 · GitHub merge queue ships;
Bors-NG deprecates itself

2025–26 · Rust rewrites bors;
satellite repos take the platform queue,
rust-lang/rust keeps its own serial lane

Each generation re-implements "test the exact merge result" and changes only the throughput mechanism; dates from repository history and project documents [graydon/bors, Tide docs, TMIB 76, Rust infra recap].
Diagram source
Scope

This guide covers merge queues for shared-branch integrity: the mechanism, the decisions, and the production failure record. It does not cover monorepo-internal systems with no public repository trail (Google's TAP, Meta's internal equivalents), CI cost optimisation as a topic of its own, or the commercial queue products. One known limit of this session: its network reached only GitHub-hosted content, so the single peer-reviewed treatment, Uber's SubmitQueue paper at EuroSys 2019, is acknowledged where relevant but never relied on, and no conference talks are cited. The evidence here is repository documents, committed postmortems and measured git history.

02

How it is actually built

Six independent implementations converge on the same five components; they diverge on exactly two questions, how candidates are constructed and who gets blamed when one fails.

Figure 2 · The reference shape, common to every implementation found

The part that varies

Intake: intent, not merge

merge commit on a
staging branch

all green

failure

Approved change
(r+, label, queue button)

Queue manager
(ordering, priorities, tree state)

Candidate constructor
single / batch / speculative chain

CI on the candidate,
never on the PR branch

Verdict watcher
(webhooks + reconciliation poll)

Fast-forward mainline
to the tested commit

Blame assigner
evict / bisect / reset / unroll

Mainline

The part that varies

Intake: intent, not merge

merge commit on a
staging branch

all green

failure

Approved change
(r+, label, queue button)

Queue manager
(ordering, priorities, tree state)

Candidate constructor
single / batch / speculative chain

CI on the candidate,
never on the PR branch

Verdict watcher
(webhooks + reconciliation poll)

Fast-forward mainline
to the tested commit

Blame assigner
evict / bisect / reset / unroll

Mainline

Every box appears in Homu, Bors-NG, the new bors, Tide, Zuul, GitLab trains and GitHub's queue; the two highlighted boxes are where implementations diverge. Reconstructed from [bors design doc], [Bors-NG README], [Zuul gating] and [GitHub docs].
Diagram source

The five common components, each attributable across the corpus:

Queue manager

Holds the approved-but-unmerged set and its order. Priorities exist everywhere because starvation is real: Homu's p=, Tide's batch-over-single rule, GitHub's queue-jump. Tree state lives here too; Rust's bot can close the whole tree, which is what "the tree had to be closed" means in the July 2026 postmortem.

Runs this way at: Rust, Kubernetes

Candidate constructor

Builds the commit that will be tested and, unchanged, published: Homu and the new bors merge onto auto-style staging branches, Bors-NG onto staging, GitHub onto temporary main/pr-N branches. Two staging branches per flavour is the recurring trick, because no provider API sets-and-merges a branch atomically [14].

Runs this way at: Rust, GitHub

Verdict watcher

Decides "CI is done and green". Harder than it sounds: the new bors abandoned its first check-suite design over webhook ordering races, and now counts completed workflows itself [14]. Homu-era repos needed a fake aggregate job just to produce one signal.

Runs this way at: Rust

Blame assigner

On failure, something must decide which change pays. Bors-NG bisects the batch; Zuul, GitLab and GitHub evict the head and rebuild everything behind it; marge-bot gives up on the batch and lands the first MR alone; the new bors runs post-merge "unrolled" builds per rollup member to attribute regressions after the fact.

Runs this way at: Bors-NG, Smarkets

Fast-forward publisher

The invariant lives here: mainline only ever moves to a commit that already passed, bit for bit. Graydon's 2013 loop ends "if ffwd works, close pull req"; the 2026 bors design ends "the base branch is fast-forwarded to the merge commit". Thirteen years, same final instruction.

Runs this way at: graydon/bors, Rust

State store (the divergence nobody advertises)

Graydon's bors was deliberately stateless, reloading everything from the API each minute; Homu added a database "because of GitHub's rate limiting"; Tide went the other way and re-derives its pool from GraphQL searches so one instance can serve dozens of orgs. Rate limits, not elegance, decided all three shapes.

Sources: Homu README, Tide docs

What the constructor builds is the real fork in the family tree. Serial construction (one candidate at a time) is what graydon/bors, Homu, marge-bot's default mode and the new Rust bors do. Batch construction (many PRs in one candidate) is Bors-NG and Tide. Chain construction (candidate N speculatively includes candidates 1..N-1, all tested in parallel) is Zuul's dependent pipeline, GitLab's train and GitHub's merge queue [1][23][22]. Chains buy wall-clock speed with compute: Zuul's own documentation states the worst case plainly, "changes are tested one at a time (as each subsequent change fails, changes behind it start again)", which means under a high failure rate the expensive parallel system quietly becomes the cheap serial one plus the bill for every discarded build.

One more architectural fact that only shows up when you read several of these at once: every component above talks to the code host through the same rate-limited, eventually consistent API that ordinary users share. Tide's entire architecture is shaped by token budgets; its maintainer guide documents GitHub's search index silently corrupting and hiding mergeable PRs from the queue until someone leaves a comment [20]. A merge queue is a distributed system whose source of truth is someone else's cache.

03

The decisions that matter

Five forks in the road, each with the recorded reason and the condition that flips it. The first one is the one teams get wrong by default.

Decision 1: take the platform's queue, or own the bot?

Chosen
  • Bors-NG's maintainer chose the platform for everyone, deprecating his own tool in May 2023: GitHub "can fix" the integration bugs a third party cannot, and gets a button where the bot gets a comment command [7]
  • Rust's satellite repos (clippy and friends) also took GitHub's queue [17]
Rejected (by Rust, for rust-lang/rust)
  • The flagship repo got a bespoke rewrite instead: try builds, human rollups, risk grades, tree closure and post-merge unrolling have no merge_group equivalent [14]
  • Even the satellite repos lost reviewer delegation on the platform queue and rebuilt it in triagebot [17]
Flips when
  • Your workflow is "approve, then land" and CI is minutes: platform queue, nothing to maintain
  • The queue is a control plane (tree state, try builds, perf runs, priorities): the platform feature is a component, not a replacement

Decision 2: serialise, batch, or speculate?

Chosen, by context
  • Serial: graydon/bors, Homu, the 2026 Rust rewrite ("only one auto build runs at a time") [14]
  • Batch + bisect: Bors-NG, turning O(N) builds into O(E log N) [5]
  • Speculative chains: Zuul, GitLab (20 parallel pipelines by default), GitHub (1–100 concurrent groups) [1][23][22]
What each gives up
  • Serial: throughput is 24h ÷ CI duration, full stop
  • Batches: a red batch delays every member until bisection finds the culprit
  • Chains: every failure discards all work behind it; compute cost scales with failure rate
Flips when
  • Arrival rate × CI duration < ~1 build-slot: stay serial, it is the simplest correct thing
  • Failures are rare and random: speculate or batch, the math favours it
  • Failures are frequent or correlated: speculation degenerates to serial at parallel prices; fix CI first

Decision 3: machine batching or human-curated batching?

Chosen by Rust
  • Humans assemble rollups from PRs the reviewers graded always / maybe / iffy / never; 92% of September 2026 PRs landed this way [11][26]
  • The procedure doc is explicit that risk-grading is judgement: CI-touching PRs are iffy, doc-only PRs always
Rejected (for rust-lang/rust)
  • Blind machine batching, which treats members as uniformly risky and pays log-factor bisection on every failure
  • Rust inverts it: even failure handling is human ("use the logs to bisect the failure to a specific PR and unapprove it") [11]
Flips when
  • No reviewer base to grade risk, or uniform small changes: machine batching wins, bisection is cheap
  • A 3–4 h CI makes every wasted build expensive and a rollup culture exists: curation wins

Figure 5 · Rust's human-curated batch: the rollup pipeline

pass

fail

Review: r+ with grade
always / maybe / iffy / never

Queue, sorted: rollups
wait at the bottom

Human packs ~10 graded
PRs into one rollup PR

Serial lane: one ~3.5 h
build of main + rollup

Fast-forward main:
10 PRs per CI cycle

Human bisects from logs,
r- the offender, repack

Post-merge unrolled builds:
each member alone,
for blame and perf

pass

fail

Review: r+ with grade
always / maybe / iffy / never

Queue, sorted: rollups
wait at the bottom

Human packs ~10 graded
PRs into one rollup PR

Serial lane: one ~3.5 h
build of main + rollup

Fast-forward main:
10 PRs per CI cycle

Human bisects from logs,
r- the offender, repack

Post-merge unrolled builds:
each member alone,
for blame and perf

Risk grading happens at review time, packing is a human act, and blame runs twice: by hand on failure, and automatically after merge via unrolled builds [rollup procedure, bors design doc].
Diagram source

Decision 4: how strict about flaky tests?

The accommodations shipped
  • Chromium's CQ retries failed shards "to work around" flake [25]
  • GitHub's queue has a setting that merges groups whose earlier members failed, "useful if you have intermittent test failures" [22]
  • Rust's procedure: @bors retry on spurious failures [11]
The strict alternative
  • Treat every red as real: GitLab restarts every pipeline behind a removed MR [23]
  • Strictness is correct and expensive; with flake it converts the queue into a random-delay generator
Flips when
  • Flake rate × batch size ≈ 1 failure per batch: accommodations stop being optional
  • But each accommodation weakens the invariant; GitHub's setting can merge a group whose head passed while a member failed

Decision 5: where does queue state live?

Chosen, by context
  • Tide: stateless, re-derived from GraphQL search each cycle, so one instance serves dozens of orgs on one token [19]
  • Homu and the new bors: a database, "essential because of GitHub's rate limiting" in 2014 and still true in 2026 [3][14]
The cost each accepted
  • Stateless-from-search inherits the provider's index bugs: corrupted search silently hides PRs from Tide [20]
  • Stateful bots must reconcile: the new bors re-syncs PR state and re-reads config every few minutes anyway [14]
Flips when
  • Many repos, one bot: search-driven statelessness is the only thing that fits the rate limit
  • One flagship repo, rich workflow: a database plus periodic reconciliation is simpler to reason about

Figure 3 · Choosing a queue in 2026

no, few committers

yes

yes

no

yes

no, risk varies
by change

no

yes

Is a broken mainline
actually expensive for you?

Protected branch +
required PR checks is enough

Full CI under ~30 min
and merges < ~20/day?

Use the platform queue
(GitHub merge queue / GitLab train)
with small groups

Are failures mostly
rare and independent?

Platform queue with
speculative groups sized by
failure rate; budget for resets

Do you need try builds,
priorities, tree closure,
risk-graded batching?

Fix CI flake first,
then revisit this tree

Own the bot: budget for a
production service with its own
postmortems and on-call

no, few committers

yes

yes

no

yes

no, risk varies
by change

no

yes

Is a broken mainline
actually expensive for you?

Protected branch +
required PR checks is enough

Full CI under ~30 min
and merges < ~20/day?

Use the platform queue
(GitHub merge queue / GitLab train)
with small groups

Are failures mostly
rare and independent?

Platform queue with
speculative groups sized by
failure rate; budget for resets

Do you need try builds,
priorities, tree closure,
risk-graded batching?

Fix CI flake first,
then revisit this tree

Own the bot: budget for a
production service with its own
postmortems and on-call

Terminal nodes are actions. The left exit is the correct default; the bottom-right exit is Rust's position and it is expensive to hold [TMIB 76, Rust infra recap].
Diagram source
DecisionChosenRejectedBecauseEvidence
Platform vs bot (ecosystem)PlatformKeep maintaining Bors-NGPlatform fixes integration seams a bot cannot reachTMIB 76, 2023
Platform vs bot (rust-lang/rust)Bespoke rewriteGitHub merge queueWorkflow depth: try, rollups, grades, tree state; delegation already lost on satellite reposInfra recap, 2026
Throughput mechanismSpeculative chainsSerial testingCI duration made serial waits "a long time"; worst case acceptedZuul gating docs
Throughput mechanism (Rust)Serial + human rollupsMachine speculationRisk is legible to reviewers; rollups carry 92% of PRs (measured)rust-forge
Batch failure handlingBisectionEvict whole batchO(E log N) builds instead of losing N green PRsBors-NG README
Batch failure handling (marge-bot)Fall back to first MRBisectionSimplicity; accepts slower recoverymarge-bot README
CI completion signalCount completed workflowsPoll check suitesCompletion webhooks arrive out of order; polling racedbors design doc
Queue stateGraphQL search, statelessPer-repo bot stateOne token must serve dozens of orgsTide docs
04

What broke in production

Four failure classes cover the public record: the queue inherits its CI's dependencies, speculation collapses under failure, the queue service itself breaks, and the platform underneath it lies.

Figure 4 · The gate reset: how one bad change erases the pipeline behind it

MainlineCI poolQueue managerMainlineCI poolQueue managerA is evicted. Builds 2 and 3contained A,so both are discarded, green ornot.Cost of A's failure: two fulldiscarded builds, plus latency for B and Cbuild 1 = main + Abuild 2 = main + A + B(speculative)build 3 = main + A + B + C(speculative)build 1 FAILEDbuild 4 = main + Bbuild 5 = main + B + C(speculative)builds 4 and 5 passfast-forward to B, then C
MainlineCI poolQueue managerMainlineCI poolQueue managerA is evicted. Builds 2 and 3contained A,so both are discarded, green ornot.Cost of A's failure: two fulldiscarded builds, plus latency for B and Cbuild 1 = main + Abuild 2 = main + A + B(speculative)build 3 = main + A + B + C(speculative)build 1 FAILEDbuild 4 = main + Bbuild 5 = main + B + C(speculative)builds 4 and 5 passfast-forward to B, then C
Builds 2 and 3 were built on the assumption A merges; A's failure makes their results unusable regardless of outcome. This is Zuul's documented worst case and the exact behaviour GitLab documents as "all pipelines behind the removed merge request restart" [Zuul, GitLab].
Diagram source
Postmortem

The queue inherited kernel.org's uptime

AssumptionCI availability is the project's own concern; external artifact hosts are background noise.
What happenedA kernel.org mirror misconfiguration made kernel source artifacts unavailable. One rust-lang/rust job fetched them directly and seventeen more fetched them transitively through crosstool-ng, so every queue build went red and the tree was closed. Four mitigation attempts over three days; the mitigation itself then broke the stable-branch artifacts PR a day later.
Blast radiusMerges to rust-lang/rust halted for roughly four days (reported 2026-07-02 07:48 UTC, tree re-opened 2026-07-06 04:26 UTC); secondary outage in stdarch.
FixMirror the artifacts in project-owned storage; a written policy that new CI jobs may not add external fetch dependencies without justification; longer-term, an allow-list.
Design ruleA merge queue multiplies CI availability into merge availability: the postmortem's own model shows five external dependencies at 99% uptime push CI below 95%. Count the external fetches in your gate jobs; that number is your queue's hidden SLA.
Postmortem

The merge bot merged a bug into itself

AssumptionThe tool enforcing "nothing lands untested" is itself protected by its tests.
What happenedBors-NG shipped a PR where an Elixir struct hit the Access protocol; every r+ crashed instead of starting a batch. The unit test used a plain map and the integration test left the field nil, so both passed while production failed.
Blast radiusThe hosted instance was down from 17:47 UTC to 02:00 UTC the next day, about eight hours, for every repository on it; a queue outage is a merge freeze for all tenants.
FixIntegration tests against the real GitHub API, acknowledged as unimplemented because they are "complicated".
Design ruleThe queue is a tier-1 production service: it needs its own postmortems, its own staging and a bypass procedure, because when it breaks the whole organisation's write path breaks with it.
Source

A failing batch starved every merge

AssumptionBatching only helps: prioritising batches over single PRs maximises throughput.
What happenedKubernetes' Tide prioritises pending batches and will hold a green single PR while a batch runs. Issue #13551, "tide: serial merges should occur when batches fail", records the consequence: a batch containing one failing PR was retested for hours while nothing merged. The thread body was unreachable from this session; the claim rests on the project's own issue title plus Tide's documented batch-priority behaviour.
Blast radiusReported as hours of zero merge throughput on affected repos; duration unverified here.
FixFall back to serial merges when batches fail, per the issue's title; Tide also bounds batches via batch_size_limit.
Design ruleAny "prefer the batch" rule needs a starvation escape hatch: a maximum hold time or an automatic fall-back to singles, or one bad PR rations everyone's merges.
Platform seam

The webhooks arrived in the wrong order

Assumption"Check suite completed" means all its workflows have reported their conclusions.
What happenedThe new Rust bors first polled check suites; GitHub would deliver a completion webhook while the API still reported the suite pending, and the suite-completed event can arrive before the last workflow-completed event. A build could be marked finished without its real conclusion.
Blast radiusDesign-level: caught and fixed before it corrupted merges; the cost was a redesign of the completion logic.
FixIgnore the suite-completed event entirely; on each workflow-completed webhook, enumerate the suite's workflows and finish only when all are complete, with a reconciliation poll as backstop.
Design ruleTreat provider webhooks as unordered hints, never as state transitions; derive state from a full read, and keep a periodic reconciliation loop for the events that never arrive. Tide's corrupted-search-index workaround is the same rule from the other direction.

The class the record does not contain is worth naming: no public postmortem in this corpus describes the invariant itself failing, a tested-and-merged commit turning mainline red. The failures are all availability failures (nothing merges) rather than integrity failures (the wrong thing merges). That asymmetry is the strongest argument in the whole record that the pattern works, and it is also exactly what you would expect if integrity failures were quietly fixed-forward without a write-up; both readings are open. The one documented near-exception is GitHub's own flake accommodation, which deliberately merges groups containing a member whose checks failed [22], trading a little integrity back for throughput.

05

Numbers you can plan against

The arithmetic of a queue is unforgiving: arrivals per day times CI hours per build must come out under 24, and every mechanism in this guide is a way of bending one of those two factors.

MetricValueAtContextAs ofSource
PRs landed through one serial queue737 / monthrust-lang/rustMeasured: first-parent git history, Sept 20262026-09git history
Successful queue builds130 / month (4.3/day)rust-lang/rustMeasured: "Auto merge of #" commits; failed attempts leave no trace in git2026-09git history
Share of PRs landing via rollup91.9%rust-lang/rustMeasured: 677 of 737 PRs rode 70 rollup builds, mean 9.7 PRs each2026-09git history
Duration of one queue build~3.5 hrust-lang/rustProject's own procedure doc; implies a hard ceiling of ~6–7 builds/day2026-09rust-forge
Queue merges379 / month (12.6/day)kubernetes/kubernetesMeasured: first-parent merge commits by kubernetes-prow[bot]2026-09git history
Builds, serial vs batched strategyO(N) vs O(E log N)Bors-NGClaimed by the project; N = PRs, E = failing PRs2023README
Parallel speculative pipelines20 defaultGitLab merge trainsVendor default limit; 1 = serial mode2026GitLab docs
Concurrent merge-group builds1–100GitHub merge queueVendor-configurable throttle on speculative CI spend2026GitHub docs
Mergeability check latency30 min → 1 minRust borsReported: switching the check from REST to GraphQL2026-07infra recap
External deps before CI < 95% uptime~5 at 99% eachRust CI modelPostmortem's own availability model (~50 deps at 99.9%, ~10 at 99.5%)2026-07postmortem
Queue-service outage~8 hbors.tech hosted instanceReported: bad deploy of the bot froze merges for all tenants2017-08retrospective
Tree closure from CI dependency outage~4 daysrust-lang/rustReported: 2026-07-02 07:48 to 2026-07-06 04:26 UTC2026-07postmortem

The planning arithmetic, worked. rust-lang/rust receives roughly 25 landable PRs a day (measured) against a serial lane that fits at most 6–7 builds of ~3.5 h (derived from the two sourced figures above; September landed 4.3 successful builds per day, the gap being failed and retried attempts). Without aggregation the queue diverges by a factor of four to five; the rollup mechanism is what closes it, packing a mean 9.7 PRs into each of 70 builds. That is the general law: a queue is viable when (arrivals × CI-hours) ÷ (PRs per build × 24) < 1, and your levers are exactly three: shorten CI, batch more PRs per build, or run speculative builds in parallel and pay for the resets.

Read these carefully

The rust-lang/rust and kubernetes/kubernetes rows are this session's own measurements from cloned git history; the method is in the ledger and anyone can re-run it. Git only records successful builds, so builds-per-day understates CI load, possibly by a lot under flake. The GitLab and GitHub rows are vendor-documented limits, not measurements. No independent published measurement of speculative-queue efficiency was reachable in this session; the one peer-reviewed source (Uber's EuroSys 2019 paper) was not fetchable, so treat any claim about speculation win-rates as unverified until you measure your own reset frequency.

06

The evidence wall

Every source behind this page, graded. All of it is repository content: this session's network reached only GitHub, which for this topic happens to be where the primary record lives. Filter by kind.

Postmortem Rust project2026-07

rust-lang/rust CI outage, 2026-07-02

A committed, full-format postmortem: kernel.org mirror outage, eighteen affected jobs (one direct fetch, seventeen transitive via crosstool-ng), four mitigation attempts, a mitigation-induced secondary failure, and an availability model for external dependencies.

Carry forwardYour queue's availability is the product of every URL your gate jobs fetch. Mirror or delete them.
raw.githubusercontent.com/rust-lang/infra-team/.../20260702-rust-outage
Postmortem bors.tech2017-08

"Downtime from 5:47 PM to 2:00 AM UTC the next day"

The hosted merge bot deployed a broken PR through its own queue; unit and integration tests both missed it for different reasons. Honest to the point of tagging itself "pants-on-head-stupid".

Carry forwardThe queue is a single point of failure for the entire write path; stage and canary it like one.
raw.githubusercontent.com/bors-ng/bors-ng.github.io/.../we-were-down
ADR bors.tech2023-05

TMIB 76: the deprecation decision

The maintainer's reasoned surrender to the platform, with the specific unfixable bugs enumerated: squash merges that cannot be marked merged, status checks that break when copied across branches, a merge button users should not see.

Carry forwardThird-party queues die on integration seams, not algorithms. List the seams before building on one.
raw.githubusercontent.com/bors-ng/bors-ng.github.io/.../tmib-76
ADR Rust project2026

New bors design document

The 2026 rewrite's architecture: one auto build at a time, two staging branches per build flavour because the API cannot merge atomically, post-merge unrolled builds for rollup blame, and a documented rejected design (check-suite polling) with the race that killed it.

Carry forwardWebhooks are hints; completion is a derived state you compute, re-check and reconcile.
raw.githubusercontent.com/rust-lang/bors/main/docs/design.md
ADR OpenStack≤2019

Zuul: Project Gating

The clearest statement of the invariant and of speculative execution's bargain, including the degenerate case. Preserved on the retired GitHub mirror; the project itself moved to OpenDev.

Carry forwardSpeculation's worst case is serial testing plus the cost of every discarded build; size it by failure rate.
raw.githubusercontent.com/openstack-infra/zuul/.../gating.rst
ADR Rust project2026-09

Rollup Procedure (rust-forge)

The operating manual for human-curated batching: the four risk grades, who may assemble a rollup, how to bisect a failed one by hand, and the queue-fairness etiquette. Contains the sentence that reframes the whole topic: the queue's job is to test PRs, not to land them.

Carry forwardIf reviewers can see risk, batching by judgement beats batching by scheduler.
raw.githubusercontent.com/rust-lang/rust-forge/main/src/release/rollups.md
ADR Kubernetes2026-06

Tide documentation and history

Created 2017 to replace mungegithub's Submit Queue; the design brief was API-token economics: identify mergeable PRs via GraphQL search so a single instance covers dozens of orgs. Batches "whenever possible".

Carry forwardRate limits are an architectural force; they decided state placement in three of the six systems here.
raw.githubusercontent.com/kubernetes-sigs/prow/.../tide/_index.md
Source Kubernetes2026

Maintainer's Guide to Tide

Operational truths: human merges invalidate the whole pool's tests, batches outrank singles, and GitHub's search index can corrupt and hide mergeable PRs until any update triggers reindexing.

Carry forwardDisable the humans: one manual merge resets every running test in the pool.
raw.githubusercontent.com/kubernetes-sigs/prow/.../tide/maintainers.md
Source Rust project / barosl2014–2025

Homu: README and eleven years of history

Why statelessness lost (rate limits), why webhooks beat polling, and the argument that build badges prove the problem exists. Its git history records its own retirement: "Disable try builds", 2025-07-22, as the replacement took over.

Carry forwardIf the default branch could never break, the status badge would be pointless; the badge is the confession.
raw.githubusercontent.com/rust-lang/homu/master/README.md
Source Mozilla / Graydon Hoare2013

The original bors

A stateless cron loop: load everything, advance the ripest PR one state, exit. The state machine in the README is the whole pattern in nine lines, ending with the fast-forward and its failure case ("someone moved master on us").

Carry forwardThe minimal correct queue is a weekend project; everything after that is throughput and platform seams.
raw.githubusercontent.com/graydon/bors/master/README.md
Source bors-ng2017–2023

Bors-NG README

The bifurcate/bifurcateCrab semantic-conflict example, the batch-then-bisect algorithm walked through on a three-PR failure, and the complexity claim O(N) vs O(E log N). Now opens with its own deprecation notice.

Carry forwardBisection makes batch cost scale with failures, not batch size; it is the right default when risk is illegible.
raw.githubusercontent.com/bors-ng/bors-ng/master/README.md
Source Rust project2026-09 (measured 2026-10-04)

rust-lang/rust git history, measured

130 successful queue builds landing 737 PRs in September 2026; 70 rollup builds carried 677 of them. First-parent commit messages make the queue's whole output auditable by anyone with a clone.

Carry forwardA queue's merge commits are a free, honest dataset; measure yours before redesigning anything.
github.com/rust-lang/rust
Source Kubernetes2026-09 (measured 2026-10-04)

kubernetes/kubernetes git history, measured

379 first-parent merge commits in September 2026, all authored by kubernetes-prow[bot]: a fully automated write path at 12.6 merges/day, batched opportunistically by Tide.

Carry forwardAt k8s scale the bot owns the merge button outright; humans express intent through labels only.
github.com/kubernetes/kubernetes
Source Kubernetes2019 (title only)

test-infra #13551: batch failure starves merges

Cited by title, flagged accordingly: "tide: serial merges should occur when batches fail". The thread body was unreachable under this session's network policy; the failure mode it names is corroborated by Tide's documented batch-priority rule.

Carry forwardEvery batch-first policy needs a documented starvation exit.
github.com/kubernetes/test-infra/issues/13551
Blog Rust project2025-10 → 2026-07

Infrastructure team quarterly recaps

The migration diary: try builds on the new bors from July 2025; rust-lang/rust merges by Q4 2025, "completing the migration off Homu"; the GraphQL mergeability win (30 min to 1 min); and the admission that satellite repos on GitHub's queue lost reviewer delegation and got it rebuilt in triagebot.

Carry forwardMigrating a queue is done in slices, lowest-stakes flavour first; try builds were the canary for a year.
raw.githubusercontent.com/rust-lang/blog.rust-lang.org/.../infrastructure-team-2026-q2-recap-and-q3-plan
Blog bors.tech2017

"About semantic conflicts" and the Whirlwind lineage guide

The pitch essay demonstrates isolation-testing's blind spot in three slides; the lineage guide records why each generation exists, including "Homu tests pull requests one at a time" as Bors-NG's founding complaint.

Carry forwardEach rewrite in this lineage was motivated by one named deficiency, not by rot; name yours before rewriting.
raw.githubusercontent.com/bors-ng/bors-ng.github.io/.../whirlwind
Blog Smarkets2017–

marge-bot README

The Not Rocket Science Rule quoted with attribution, the argument that manual rebase-and-retry collapses at 5–10 minute CI with a busy team, and a third blame policy: on batch failure, land the first MR and retry the rest.

Carry forwardAt GitLab shops the same pattern re-evolved independently; the rule is platform-agnostic even when the bots are not.
raw.githubusercontent.com/smarkets/marge-bot/master/README.md
Vendor GitHub2026

Managing a merge queue (docs source)

The platform queue's mechanics from the docs repository: merge groups on temporary branches, 1–100 build concurrency, min/max group sizes with a wait timer for deploy-coupled branches, queue-jumping that rebuilds everything behind, and the flake-tolerance checkbox.

Carry forwardThe platform queue is Zuul's algorithm productised; what it lacks is everything around the algorithm.
raw.githubusercontent.com/github/docs/.../managing-a-merge-queue.md
Vendor GitLab2026

Merge trains (docs source)

Parallel merged-results pipelines, the restart cascade on failure, the 20-pipeline default limit, and "merge immediately" documented as an escape hatch that aborts the train and may leave the target branch needing "additional work".

Carry forwardEvery queue grows an admin bypass; the vendor documenting its blast radius is a gift, read it.
raw.githubusercontent.com/gitlabhq/gitlabhq/.../merge_trains.md
Vendor Google / Chromium2026

Chromium commit queue docs

A Gerrit-world data point: dry runs separated from submitting runs, automatic retry of failed shards to absorb flake, and an opt-in "Mega-CQ" that trades much longer wall time for broader coverage on risky changes.

Carry forwardRisk-tiered verification (normal vs mega) is the machine version of Rust's human risk grades.
raw.githubusercontent.com/chromium/chromium/main/docs/infra/cq.md
07

Build a miniature, then productionise it

Six rungs from reproducing the failure to operating the queue as a service. The line from toy to real is crossed at rung four.

Break main on purpose

Two PRs on a toy repo: one renames a function everywhere, one adds a call to the old name. Get both green, merge both through the normal button.

Done when: main is red while both PRs show green checks.  Teaches: isolation testing's blind spot; you now believe the pitch essay.

The not-rocket-science loop, serial

A script triggered by an "approved" label: merge the PR onto a staging branch, wait for CI on that branch, fast-forward main on green, comment on red. Disable direct pushes. This is graydon/bors in ~200 lines.

Done when: nothing reaches main except by fast-forward from a tested commit, for a week.  Teaches: the invariant, and the eventual-consistency pain of "is CI done?".

Batch and bisect

Merge all approved PRs into one candidate. On failure, split in half and requeue both halves, Bors-NG style. Inject a deliberately bad PR among nine good ones.

Done when: the bad PR is isolated and rejected in log-factor builds while the nine land.  Teaches: why batch cost scales with failures, and what the offender's author experiences.

Speculate, then measure the reset

Build candidates in a chain (B's build includes A) with parallel CI. Now fail A and watch builds behind it become garbage. Count discarded builds per injected failure at chain lengths 3, 5, 10.

Done when: you can state your reset cost as a function of chain length and failure rate.  Teaches: speculation is borrowed certainty, repaid with interest on every miss.

Inject flake

Make one test fail 2% of the time. Watch throughput fall and spurious evictions anger your imaginary contributors. Add a retry policy, then a quarantine list, and decide what the queue does with a quarantined test's verdict.

Done when: throughput recovers near baseline and every accommodation you added is written down with its integrity cost.  Teaches: the queue is a flake amplifier; Chromium's shard retries and GitHub's tolerance checkbox stop looking lazy.

Operate it as a service

Dashboards: queue depth, time-from-approval-to-merge, builds per landed PR, reset count. Add a tree-closure switch, a priority flag, and a documented human bypass with an audit trail. Then kill your bot mid-build and verify recovery.

Done when: "why isn't my PR merged" is answerable from the dashboard, and the bot's own deploy goes through a queue.  Teaches: what bors.tech learned in August 2017, before production teaches it to you.

08

Keep hunting

The queries and commands that produced this page. The unusual ones are the git commands: for this topic the primary record is repository content, and it stays current after every blog post goes stale.

Measure any project's queue from git

  • git clone --bare --filter=blob:none --shallow-since=2026-08-15 https://github.com/ORG/REPO
  • git log --first-parent --since=2026-09-01 --until=2026-10-01 --format="%s %an"
  • git log --first-parent --format=%s | grep -c "Auto merge of #"
  • git log --format=%s | grep -c "Rollup merge of #"

Find the decision record in the repo, not the blog

  • path:docs design.md merge queue
  • repo:kubernetes-sigs/prow path:site tide
  • "post-mortems" OR "retrospective" path:service-catalog
  • site:github.com inurl:issues tide batch merge starve

Failure vocabulary of this domain

  • "gate reset" OR "tree closed" OR "merge freeze" postmortem
  • "bors retry" OR "spurious failure" site:github.com
  • merge train restart flaky "intermittent test failures"
  • "merge_group" event required checks stuck

The era documents, by exact phrase

  • "not rocket science rule" bors
  • "keeping master green" submitqueue eurosys
  • "submit queue" mungegithub tide replace
  • bors-ng deprecated "merge queue" TMIB
09

References

  1. Zuul, "Project Gating" (v3 documentation) OpenStack Infra, content as of the GitHub mirror's final sync, 2019-04-16. Checked 2026-10-04.
  2. Graydon Hoare, bors (original) README GitHub, first commit 2013-02-01; maintenance note 2021. Checked 2026-10-04.
  3. Homu README GitHub (rust-lang/homu), repository active 2014-12-18 to 2025-11-04. Checked 2026-10-04.
  4. Homu: "Disable try builds" merge commit GitHub, 2025-07-22. Checked via git clone 2026-10-04; the HTML page may require sign-in.
  5. Bors-NG README GitHub (bors-ng/bors-ng), deprecation notice 2023. Checked 2026-10-04.
  6. bors.tech, "About semantic conflicts" bors-ng.github.io source, 2017-02-02. Checked 2026-10-04.
  7. bors.tech, TMIB 76: "This April, bors-ng is deprecated" bors-ng.github.io source, 2023-05-01. Checked 2026-10-04.
  8. bors.tech, TMIB 78 (final newsletter) bors-ng.github.io source, 2023-07-01. Checked 2026-10-04.
  9. bors.tech, "Whirlwind" lineage guide bors-ng.github.io source, 2017-04-06. Checked 2026-10-04.
  10. bors.tech, "Downtime from 5:47 PM to 2:00 AM UTC the next day" bors-ng.github.io source, 2017-08-24. Checked 2026-10-04.
  11. Rust Forge, "Rollup Procedure" rust-lang/rust-forge, last modified 2026-09-25. Checked 2026-10-04.
  12. Rust Forge, "Bors" service page rust-lang/rust-forge, current. Checked 2026-10-04.
  13. rust-lang/bors README GitHub, repository started 2022-11-14. Checked 2026-10-04.
  14. rust-lang/bors design document GitHub, current as of 2026-10-02 (last repo commit). Checked 2026-10-04.
  15. Rust Infrastructure Team, 2025 Q3 recap Inside Rust blog source, 2025-10-16. Checked 2026-10-04.
  16. Rust Infrastructure Team, 2025 Q4 recap Inside Rust blog source, 2026-01-13. Checked 2026-10-04.
  17. Rust Infrastructure Team, 2026 Q2 recap Inside Rust blog source, 2026-07-15. Checked 2026-10-04.
  18. Rust Infrastructure Team, "2026-07-02 rust-lang/rust CI outage postmortem" rust-lang/infra-team, July 2026. Checked 2026-10-04.
  19. Prow documentation, "Tide" kubernetes-sigs/prow, page last modified 2026-06-02. Checked 2026-10-04.
  20. Prow documentation, "Maintainer's Guide to Tide" kubernetes-sigs/prow, current. Checked 2026-10-04.
  21. Prow source, pkg/config/tide.go (batch_size_limit) kubernetes-sigs/prow, current. Checked 2026-10-04.
  22. GitHub Docs source, "Managing a merge queue" github/docs, current. Checked 2026-10-04.
  23. GitLab documentation source, "Merge trains" gitlabhq/gitlabhq, current. Checked 2026-10-04.
  24. Smarkets, marge-bot README GitHub, project started 2017. Checked 2026-10-04. Quotes Graydon Hoare's Not Rocket Science Rule.
  25. Chromium documentation, "Chromium Commit Queue" chromium/chromium, current. Checked 2026-10-04.
  26. rust-lang/rust repository (git history measured in this session) Measured 2026-10-04 over 2026-09-01..2026-10-01; method in the evidence ledger (sources.md).
  27. kubernetes/kubernetes repository (git history measured in this session) Measured 2026-10-04 over 2026-09-01..2026-10-01; method in the evidence ledger (sources.md).
  28. kubernetes/test-infra issue #13551, "tide: serial merges should occur when batches fail" Cited by title only; thread body unreachable under this session's network policy. Checked 2026-10-04.

Unreachable but acknowledged: Ananthanarayanan et al., "Keeping Master Green at Scale" (EuroSys 2019, Uber's SubmitQueue), Graydon Hoare's original "Not Rocket Science Rule" post (graydon2.dreamwidth.org, quoted in [24]), GitHub's merge-queue changelog entries, and matklad's 2023 merge-queue essay. This session's network reached GitHub-hosted content only; nothing above depends on the unreachable items.