Evacuating the region  / field guide
Practitioner field guide · 2026-10-11

Evacuating the region: who can actually leave, and what it costs

When the place a system runs becomes unusable, someone has to serve every user from the places that remain, using machinery that must work during the worst hour the organisation has had all year. This guide reconstructs how that move is actually made, from GitLab's public recovery runbooks and gameday measurements, Meta's and Netflix's drain practice, and the postmortems of GitHub, Cloudflare, AWS, Roblox and Datadog. After reading it you can name your own organisation's posture, price the move to the next one, and defend the number you write in the recovery-time box.

24 primary documents 10 production systems 6 incidents + 1 instructive drill Evidence through October 2026 Read: ~22 min
01

The territory

The public record splits organisations into three postures, and the honest ones are the two extremes: days, or minutes. The disasters cluster in the believed-but-unexercised middle.

96 h
Downtime pre-printed in GitLab's regional-recovery change template, Criticality 1
<10 min
To drain most of Facebook's web services out of a datacenter, without user-visible impact
+46%
Average write-latency increase Datadog measured to make PostgreSQL failover safe (synchronous commit)
24 h 11 m
GitHub's degradation after a 43-second partition triggered an automated cross-country failover

State the problem without naming a technology: your system runs in a handful of physical places, one of those places stops being usable, and every user must now be served from the places that remain. Three things have to happen at once. Traffic has to be steered away from the dead place by machinery that does not itself live there. The data the dead place was writing has to exist somewhere else, complete enough to serve from. And somebody, or something, has to decide to go, under pressure, with partial information, knowing the move itself can cause damage.

The surprise in this record is not that evacuation is hard. It is how honestly the extremes have published their numbers, and how far apart they are. GitLab's regional-recovery change template, the actual form an engineer would file while GCP's us-east1 burns, carries a pre-printed downtime of 96 hours [7]. Meta's Maelstrom paper reports draining most web services from a datacenter in under ten minutes, and that the tool ran more than 100 real mitigations in four years [21]. Both organisations are telling the truth. They have simply bought different products: one rebuilds from backups, the other treats evacuation as a routine traffic operation it performs weekly. The organisations that got hurt worst in the postmortems below, GitHub in 2018 and Cloudflare in 2023, sat between those postures: a failover believed to exist, exercised rarely or never, that behaved differently from the belief on the day it ran.

Scope, stated up front. This guide covers moving a stateful web-scale system off a failing region or datacenter: the steering, the data, the decision, and the drills. It covers zonal evacuation where the public record uses it as the rehearsal for regional loss, which is how GitLab treats it. It deliberately does not cover the internal design of multi-writer databases (Spanner and CockroachDB internals get one paragraph as a price tag, not a chapter), DNS and anycast mechanics, compliance-driven paper DR, or recovery from data corruption, which is a restore problem rather than an evacuation problem and is covered by an earlier guide in this collection.

Figure 1 · Three postures, three clocks

pay once per disaster

pay every day

Traffic shifters: clock reads minutes

All regions serve daily

Drain is a routine operation
Meta: under 10 min,
100+ real uses

Failover-ready: clock reads seconds, then hours

Warm standby region

Promote on failure detection
GitHub 2018: 43s to promote,
24h11m to recover

Rebuilders: clock reads days

Backups in
multi-region storage

Rebuild fleet in new region
GitLab target: 96h then 48h

pay once per disaster

pay every day

Traffic shifters: clock reads minutes

All regions serve daily

Drain is a routine operation
Meta: under 10 min,
100+ real uses

Failover-ready: clock reads seconds, then hours

Warm standby region

Promote on failure detection
GitHub 2018: 43s to promote,
24h11m to recover

Rebuilders: clock reads days

Backups in
multi-region storage

Rebuild fleet in new region
GitLab target: 96h then 48h

Every organisation in this corpus sits in one of three postures, and the time to evacuate is a property of the posture, not of the incident. Sources: GitLab DR blueprint, GitHub 2018, Maelstrom OSDI '18, Netflix 2016.
Diagram source
02

How it is actually built

Across GitLab, Meta, Netflix, Datadog and the failure accounts, the same five components recur. The posture an organisation holds is set by which of the five it has actually built, not by which it has documented.

Figure 2 · The reference shape of an evacuable system

Region B (surviving)

Region A (failing)

replication, lag โ‰ค budget

executes the move

shiftable fractions

shiftable fractions

Users

Steering layer
geo-DNS override / edge routing tables

Stateless serving

Primary data

Stateless serving +
pre-bought headroom

Replica / backups
with declared loss budget

Ops plane: runbooks, change issues,
monitoring, credentials

Decision authority
human change issue or drilled automation

Region B (surviving)

Region A (failing)

replication, lag โ‰ค budget

executes the move

shiftable fractions

shiftable fractions

Users

Steering layer
geo-DNS override / edge routing tables

Stateless serving

Primary data

Stateless serving +
pre-bought headroom

Replica / backups
with declared loss budget

Ops plane: runbooks, change issues,
monitoring, credentials

Decision authority
human change issue or drilled automation

The two components most often missing in the failure accounts are drawn with heavy borders: the steering layer outside any region, and the operations plane that shares no fate with the region it evacuates. Reconstructed from GitLab's recovery runbook, Taiji (SOSP '19) and Cloudflare's 2023 postmortem.
Diagram source

The five components, attributed

A steering layer that outlives the region. Netflix built tools to override geo-DNS and send all users to a healthy region, and used them in anger; its 2016 account describes a real event where traffic stayed failed over to the west region for more than 24 hours before shifting back gradually [19, 20]. Meta went further and made steering continuous: Taiji recomputes the fraction of edge traffic each datacenter receives as an optimisation problem, so a drain is a change of coefficients rather than an emergency switch [22]. The counter-example is GitLab's own blueprint, which lists, with unusual honesty, "we have not standardized a way to divert traffic at the edge from 1 region to another" among the binding constraints on its recovery time [4]. If the steering layer does not exist in advance, it gets built during the incident, and that is where the days go.

Data somewhere else, with a declared loss budget. Every evacuable design states, in configuration rather than in a policy document, how many committed writes it is prepared to lose. Patroni's documentation is the plainest statement in the corpus: in asynchronous mode "the cluster is allowed to lose some committed transactions to ensure availability", bounded by maximum_lag_on_failover bytes of transaction log plus a window of recent writes [15]. GitLab's budget is the age of a disk snapshot in a multi-region bucket, measured at about two hours of replica catch-up from a one-hour-old snapshot [6]. Datadog decided after a bad gameday that for failover candidates the budget must be zero, and paid 46% average write latency for it [12]. CockroachDB prices the zero-budget posture at the engine level: surviving a region without any failover event needs at least three regions, replication factor five, and a write-latency floor of one round trip to the nearest other region [17].

A decision authority. GitLab's answer is a human filing a Criticality 1 change issue from a template; recovery "starts with a change issue", and the gameday runbook assigns a named change technician and reviewer [6, 9]. GitHub's 2018 answer was automation: Orchestrator promoted new primaries within 43 seconds, and the postmortem's remediations included preventing it from promoting across region boundaries [13]. The tooling itself encodes the ambivalence: Orchestrator ships an anti-flapping block that refuses a second automated recovery on the same cluster until a human acknowledges the first [16]. Meta resolves the ambivalence by rehearsal rather than by hierarchy: drains are executed from runbooks by Maelstrom, and the same tool runs the weekly drain tests, so the emergency operation and the routine operation are the same code path [21].

Capacity already bought at the destination. The surviving regions must absorb the failed region's load at the worst moment. Meta added an explicit disaster-recovery buffer so healthy regions can take the shifted load [23]. Netflix shifted traffic back "more gradually than normal emergency practice" to let services scale and caches warm [19]. GitLab's blueprint again names the gap: "we don't have confidence that Google can provide us with the capacity we need in a new region, specifically the large amount of SSD necessary to restore all of our customer Git data", which is why its remediation is a pre-allocated "regional bulkhead" with quotas set and synced in the recovery region [4, 5]. Datadog's 2023 recovery found the destination limit the hard way: replacing tens of thousands of nodes at once "created a thundering herd that tested the cloud provider's regional rate limits in various ways, none of which was obvious ex ante" [10].

An operations plane that shares no fate with the region. This is the component the incidents keep re-teaching. GitLab keeps a separate ops instance precisely so that the change issue that declares the recovery can be created when GitLab.com itself is down [6]. Cloudflare's control plane was the casualty rather than the instrument in November 2023 [14]. Roblox's monitoring depended on the same Consul cluster whose failure it needed to diagnose, which the company's own report identifies as a reason 73 hours passed before full return to service [18]. And AWS's October 2025 event showed the industry-scale version: services far beyond DynamoDB stalled because control-plane dependencies concentrate in us-east-1 [11].

The divergence that defines the posture

Every system in this corpus has component two in some form; backups are universal. The posture is set by components one and five. An organisation with a steering layer and an independent ops plane can evacuate in minutes and can afford to practise; an organisation without them is, structurally, a rebuilder, whatever its DR document says. Check those two components first in any review.

03

The decisions that matter

Five forks, each with the condition that flips it. The first one prices the other four.

Decision 1 · Which posture do you buy: rebuild, standby failover, or routine shifting?

Chosen, by whom
  • GitLab: rebuild from multi-region backups, target 96h then 48h RTO, because GitLab.com runs in a single region and a second full region is a daily cost its blueprint never argues for [1]
  • Netflix and Meta: routine shifting, bought after real regional pain (the Christmas Eve 2012 ELB outage for Netflix) [20]
Rejected, stated reason
  • GitLab rejected (by omission and by scope) active-active serving; the blueprint limits DR to primary services and proposes a bulkhead, not a second live region [1, 5]
  • Netflix rejected single-region-plus-backups after measuring what a regional ELB failure cost it in 2012 [20]
Flips when
  • Expected cost of a day of downtime, times the realistic frequency of regional loss, exceeds the daily tax of a second live region (capacity, replication egress, engineering). Netflix crossed that line with one incident; most B2B systems with tolerant customers never do. GitLab wrote 96 hours on the form and its customers stayed.

Decision 2 · Who pulls the trigger between regions: automation or a human?

Chosen, by whom
  • GitHub after 2018: humans; a remediation was preventing Orchestrator from promoting primaries across region boundaries [13]
  • GitLab: a human files the C1 change issue; automation executes steps inside it [6]
Rejected, stated reason
  • Unrestricted cross-region automation: a 43-second partition met rules that were locally correct and produced a promotion no operator would have ordered [13]
  • Fully manual recovery inside a zone: Patroni and Orchestrator both automate that layer and the record shows no regrets there
Flips when
  • The operation is exercised often enough that running it is unremarkable. Meta's drains are automated because the same automation runs weekly as a test [21]. The rule the record supports: automate within the blast radius you drill at that cadence; a human decides at the boundary you drill rarely.

Decision 3 · Take the loss budget, or pay the write tax?

Chosen, by whom
  • Datadog, 2026: synchronous replication for failover candidates, measured at +46% average write latency (synchronous_commit=on), after a gameday left clusters with no safe candidate [12]
  • Patroni's default posture: async with a bounded budget, "allowed to lose some committed transactions to ensure availability" [15]
Rejected, stated reason
  • Datadog rejected DRBD block-level replication after evaluation, and has not adopted quorum commit ("we did not test synchronous replication with quorum commit mode") [12]
  • CockroachDB's zero-failover mode is priced, not rejected: at least 3 regions, replication factor 5, writes slower by one inter-region round trip [17]
Flips when
  • The application cannot reconcile forked writes after the fact. GitHub could not, and spent the balance of 24 hours restoring consistency [13]. If reconciliation is impossible or manual, the write tax is cheaper than it looks; if writes are idempotent events a replay can rebuild, the loss budget is honest and much cheaper.

Decision 4 · Where does the operations plane live?

Chosen, by whom
  • GitLab: a separate ops instance, named in the runbook as the place to declare the change when GitLab.com is unavailable [6]
  • Cloudflare after November 2023: stand every product up on the DR site, since "a handful of products did not properly get stood up" the first time [14]
Rejected, stated reason
  • Running monitoring and recovery tooling inside the thing being recovered: Roblox names its telemetry's dependence on Consul as a driver of the 73-hour diagnosis [18]
  • GitLab's own Chef server remains a single-zone dependency with a 4-hour snapshot RPO, acknowledged in the blueprint rather than fixed [1]
Flips when
  • Never, on the evidence here. No account in this corpus regrets separating the ops plane. The real decision is how much of it to separate; GitLab's answer is the change system, runbooks, and monitoring reads, which is the minimum that lets a recovery start.

Decision 5 · After the region returns: fail back now, or stay?

Chosen, by whom
  • GitHub 2018: stay on the West Coast despite cross-country latency, because the East Coast held unreplicated writes and failing back risked them [13]
  • Netflix: return gradually, hours not seconds, to let caches warm [19]
Rejected, stated reason
  • Immediate failback: GitLab's runbook warns that components like Gitaly "may incur more downtime when falling back to the old zone" [6]
Flips when
  • Data diverged during the event. If both sides accepted writes, failback is a reconciliation project, not a traffic operation; if the evacuation was clean, failback is the same drill run in reverse, at cache-warming speed.

Figure 3 · Choosing a posture you can defend

yes

no

yes

no

yes

no: believed-but-unexercised
standby, where the 2018 and 2023
disasters live

Can the business absorb
days of downtime after
a true regional loss?

Be a rebuilder, deliberately:
multi-region backups, bulkhead quotas,
ops plane elsewhere, write 96h on the form

Can the product tolerate
losing a bounded window
of committed writes?

Will you drain traffic
in production at least
quarterly?

Pay the write tax:
sync replication or 3-region quorum,
then drill the steering layer

Traffic shifter:
steering layer + capacity buffer,
drain as routine operation

Re-ask question 1
honestly

yes

no

yes

no

yes

no: believed-but-unexercised
standby, where the 2018 and 2023
disasters live

Can the business absorb
days of downtime after
a true regional loss?

Be a rebuilder, deliberately:
multi-region backups, bulkhead quotas,
ops plane elsewhere, write 96h on the form

Can the product tolerate
losing a bounded window
of committed writes?

Will you drain traffic
in production at least
quarterly?

Pay the write tax:
sync replication or 3-region quorum,
then drill the steering layer

Traffic shifter:
steering layer + capacity buffer,
drain as routine operation

Re-ask question 1
honestly

Terminal nodes are commitments, and each carries its price from section 5. The dangerous exit is the dashed one: a standby you believe in but do not drill is the GitHub-2018 and Cloudflare-2023 position. Derived from decisions 1 to 3 above.
Diagram source
04

What broke in production

Three failure classes account for every incident in this corpus: the failover that ran when it should not have, the failover that could not run, and the failure bigger than any region.

Class one · the failover that ran

Postmortem

GitHub, October 2018: 43 seconds of partition, 24 hours of repair

AssumptionAutomated primary promotion is safe wherever a quorum of the failover tooling can form one.
What happenedRoutine maintenance partitioned the East Coast hub from its datacenter for 43 seconds. Orchestrator nodes on the West Coast and in the cloud formed a quorum and promoted West Coast primaries. When the partition healed seconds later, the East Coast held writes the West had never seen, and applications were already writing to the West.
Blast radius24 hours 11 minutes of degraded service; webhooks and Pages paused; days of background reconciliation of the forked writes. No user data lost, by GitHub's account.
FixAmong the published follow-ups: prevent Orchestrator from promoting primaries across region boundaries, and institutionalise the status mechanism that kept integrity ahead of usability.
Design ruleScope automated failover to the boundary where promotion is cheap to undo. Across a boundary where undo means reconciling forked writes, the trigger belongs to a human, however good the detector is.

Figure 4 · The 2018 failure path, step by step

"West replicas""East primaries""Orchestrator quorum(West + cloud)""Network""West replicas""East primaries""Orchestrator quorum(West + cloud)""Network"holds writes neverreplicated Westtwo primaries' histories haveforked24h11m: restore from backups,replay,reconcile22:52 UTC partition (43s)detect East loss, promote West to primarypartition healsEast demoted, cannot rejoincleanlyapplications nowwriting Westhumans halt furtherfailover
"West replicas""East primaries""Orchestrator quorum(West + cloud)""Network""West replicas""East primaries""Orchestrator quorum(West + cloud)""Network"holds writes neverreplicated Westtwo primaries' histories haveforked24h11m: restore from backups,replay,reconcile22:52 UTC partition (43s)detect East loss, promote West to primarypartition healsEast demoted, cannot rejoincleanlyapplications nowwriting Westhumans halt furtherfailover
The promotion was correct by the tooling's rules at every step; the damage came from writes accepted on both sides of a healed partition. Reconstructed from GitHub's post-incident analysis.
Diagram source

Class two · the failover that could not run

Gameday finding

Datadog, 2026: every candidate too stale to promote

AssumptionIf the primary fails, some replica can always be promoted.
What happenedA gameday injected network latency into one zone. Primaries kept accepting writes while replication lag grew; when failover was attempted, every standby exceeded maximum_lag_on_failover, and Patroni, correctly, refused to promote. The clusters were writable nowhere safe and stuck.
Blast radiusA drill, not an outage, which is the point: the failure was discovered at rehearsal price. The only options live were waiting out the latency or accepting data loss.
FixSynchronous replication for failover candidates, tuned with Patroni; read replicas stay async to contain the cost. Measured price: 32% to 53% higher average write latency depending on commit mode.
Design ruleAn async replica is not a failover target during the failures that matter, because the same network fault that kills the primary inflates the lag that disqualifies the replica. Decide at design time whether the lag guard will ever say no, and what the answer is when it does.
Postmortem

Cloudflare, November 2023: the DR site held only what had been put on it

AssumptionHigh availability across the three core datacenters would carry the control plane through the loss of one of them.
What happenedA total power failure took out the primary control-plane facility. Core services came up at the disaster recovery site within roughly six hours, but, in the postmortem's words, "a handful of products did not properly get stood up on our disaster recovery sites. These tended to be newer products where we had not fully implemented and tested a disaster recovery procedure."
Blast radiusControl plane and analytics degraded from 2023-11-02 11:43 UTC to 2023-11-04 04:25 UTC; log processing waited for the original facility to return. The data plane kept serving throughout.
FixPublished commitments to put every product on the HA and DR footing and to test it; the following March, the same facility failed again and APIs and dashboards recovered in about seven minutes without human intervention, by Cloudflare's follow-up account.
Design ruleA DR site is an inventory, not a property. Its contents are exactly the services someone enrolled and drilled, so the review question is never "do we have DR" but "which services joined after the last test".
Postmortem

Roblox, October 2021: nowhere to go, and blind

AssumptionOne well-provisioned coordination cluster can carry every workload, and the monitoring will show what is wrong with it.
What happenedA newly enabled Consul streaming feature and a BoltDB issue degraded the single Consul cluster under load. Every backend service depended on it, so there was no healthy place to shift anything, and the telemetry needed for diagnosis depended on the same cluster.
Blast radius73 hours of outage, by Roblox's own report. Hardware swaps and a snapshot restore both failed to stick before the root cause was found.
FixPublished remediations include multiple availability zones and datacenters, removing the circular dependency in observability, and splitting workloads off the single cluster.
Design ruleEvacuation presumes a destination. A singleton coordination layer makes the whole estate one region in the sense that matters, whatever the hardware map says, and monitoring that lives inside it disappears exactly when you need it.

Figure 5 · How a cluster gets stuck: the lag guard's refusal

async replication,
lag under budget

network latency grows

primary still accepts writes,
lag climbs past budget

primary isolated or down

a standby is
within budget

every standby over
budget, guard refuses

operator accepts
bounded loss

network recovers,
replicas catch up

Healthy

Degraded

PrimaryLost

Promoted

Stuck

async replication,
lag under budget

network latency grows

primary still accepts writes,
lag climbs past budget

primary isolated or down

a standby is
within budget

every standby over
budget, guard refuses

operator accepts
bounded loss

network recovers,
replicas catch up

Healthy

Degraded

PrimaryLost

Promoted

Stuck

The guard is doing its job: it converts silent data loss into a visible, waiting state. The design question is what the operator is allowed to do from the stuck state. From Datadog's account and Patroni's replication-modes doc.
Diagram source

Class three · the failure bigger than the region

Postmortem

Datadog, March 2023: five regions, three clouds, one 06:00 UTC window

AssumptionRunning US1, EU1, US3, US4 and US5 on separate clouds and regions means no single event reaches all of them.
What happenedAn automatically applied systemd security update restarted systemd-networkd across the fleet between 06:00 and 07:00 UTC, deleting the CNI-managed routes on tens of thousands of nodes. The postmortem asks its own best question: "why did the update automatically apply in five distinct regions that span dozens of availability zones and run on three different cloud providers?"
Blast radiusAll regions, all services, from 06:03 UTC on March 8; detection in three minutes; full resolution, including backfill, on March 10 at 06:25 UTC. Mass node replacement then tripped the cloud provider's regional rate limits.
FixThe legacy auto-update channel was disabled across all regions; recovery ordering and capacity handling were reworked per the follow-up posts.
Design ruleRegions partition your request path, not your change path. Anything that applies everywhere on a clock, base images, agent updates, unattended upgrades, is a global region, and evacuation has no destination when it fires. Stagger it or it will synchronise your failures.
Postmortem

AWS, October 2025: the region everyone else's control plane lives in

AssumptionCustomers in other regions, and multi-region customers, are insulated from a us-east-1 service failure.
What happenedA latent race between DynamoDB's redundant DNS "Enactor" processes left the regional endpoint with an empty DNS record. AWS's summary describes three distinct periods of impact; the curated postmortem ledger's reading of it records the cascade into EC2's workflow manager ("congestive collapse"), Network Load Balancer health-check flapping, and Lambda, for roughly 15 hours.
Blast radiusBeginning 2025-10-19 23:48 PT; DynamoDB global tables kept serving in other regions but replication into us-east-1 lagged for hours. Dependent services and customers worldwide felt a nominally regional event.
FixAWS's published summary covers the DNS automation race and recovery throttling; the structural lesson for customers is in the dependency direction, not in AWS's patch.
Design ruleBefore pricing your own evacuation, enumerate which of your recovery steps call a control plane homed in the region you are leaving: IAM changes, DNS updates, capacity launches. A plan whose first step needs the dying region is a plan to wait.
The absence worth noticing

No postmortem in this corpus describes a planned regional evacuation that ran and failed. The failures are automated failovers that fired wrongly, failovers that could not run, and events bigger than a region. Two readings are possible: deliberate evacuations mostly work (Netflix and Meta publish only successes, and Meta reports 100+ real mitigations), or organisations that attempt and botch one do not write it up. Either way, the public record cannot tell you your evacuation works. Only your own drill can, which is why section 7 ends in production.

05

Numbers you can plan against

Every figure is dated and sourced; the note below separates measured from claimed from derived.

MetricValueAtContextAs ofSource
Drain most web services from a datacenter< 10 minMetaRoutine and emergency drains via Maelstrom; 100+ real uses over 4 years2018OSDI '18
Drain during a real fiber cut~17 min / ~1.5 hMetaMost user-facing traffic / all traffic, one incident reported in the paper2018OSDI '18
Time failed over to one region> 24 hNetflixReal event; traffic held in the west region, then returned gradually to warm caches2016Netflix tech blog
Regional recovery target (RTO / RPO)96 h / 2 h → 48 h / 0GitLabFY24 target, then FY25 target; rebuild from multi-region backups, never validated end to end at time of writing2024DR blueprint
Zonal gameday, production Gitaly restore2 h 05 mGitLabEnd-to-end drill process in production, 2024-10-21; 45 VMs provisioned in 39 min2024recovery measurements
Zonal gameday, staging, full process1 h 15 m to 4 h 15 mGitLabSeven measured Gitaly drills, 2024 to 2026, including MR creation and verification2026recovery measurements
Replica catch-up after zonal loss~2 hGitLabNew Postgres replica from a 1-hour-old disk snapshot catching up to primary2022recovery runbook
Write-latency price of safe failover+32% to +53%DatadogAverage write latency by synchronous_commit mode (local / remote_write / on / remote_apply); +46% for the mode they treat as "on"2026Datadog engineering
Price of surviving a region with no failover event≥ 3 regions, RF 5, +1 RTT writesCockroachDBDocumented requirements of SURVIVE REGION FAILURE; read latency unaffected2025CRDB docs
Partition that triggered cross-region promotion43 sGitHubFollowed by 24 h 11 m of degradation and days of reconciliation2018post-incident analysis
Control plane onto DR site~6 hCloudflare11:43 UTC failure to 17:57 UTC mostly serving from DR; log pipeline waited ~40 h for the facility2023post mortem
Outage with no evacuation destination73 hRobloxSingle Consul cluster, telemetry dependent on it2021Return to Service
Detection of a global trigger3 minDatadogInternal monitoring alert after the first faulty upgrade, 2023-03-08; full resolution ~48 h later with backfill2023deep dive
Read these carefully

Measured: the GitLab gameday rows (timestamped drill records), Datadog's latency percentages (their own benchmark) and incident timings. Reported but not independently verifiable: Meta's drain times and Netflix's event, which are the operators' own accounts of their own tooling. Documented price, not measurement: the CockroachDB row. Derived: nothing in this table; where this guide computes anything it shows the arithmetic in place. The GitLab RTO targets are targets; the blueprint itself says the regional path had not been validated end to end, so treat 96 hours as a floor for a first-time rebuild, not an estimate of one.

The cost model worth carrying: a rebuilder pays almost nothing daily (multi-region bucket storage and quota paperwork) and pays days once; a traffic shifter pays for permanent spare capacity in every surviving region (Meta's DR buffer; the headroom Netflix scales into during a shift) plus the engineering to keep the drill routine, and pays minutes once. The crossover variable is almost always the cost of a lost day, not the infrastructure bill. What nobody has published, in this corpus or anywhere this hunt reached, is the annual fully loaded cost of either posture at a named scale; that number remains open.

06

The evidence wall

Every source behind this page, graded. Cards marked with a dagger (†) sit on hosts this research environment could not open directly; their content was read through search-engine retrieval, quotes are limited to what retrieval returned verbatim, and the ledger (sources.md, shipped beside this page) records the detail.

ADRGitLab2024-01

Disaster Recovery blueprint

The working design document for GitLab.com DR: scope limited to primary services, FY24 regional targets of 96 h RTO / 2 h RPO moving to 48 h / 0, and the admission that regional recovery "has not yet been validated end-to-end, so we don't know how long the RTO is".

Carry forwardWrite the targets you can defend, then publish the gap between target and validation; the blueprint's honesty is what makes the rest of it credible.
gitlab.com/gitlab-org/gitlab · blueprints/disaster_recovery
ADRGitLab2024-01

Regional recovery proposal (the bulkhead)

Lists eleven concrete constraints on regional RTO, including no standard way to divert edge traffic, ops tooling in a single region, and no confidence the cloud can supply recovery capacity; proposes a pre-allocated "regional bulkhead" with synced quotas.

Carry forwardThe eleven-item list is a ready-made audit: run it against your own estate and the items you cannot answer are your real RTO.
gitlab.com · disaster_recovery/regional.md
ADRGitLab2024-02

Issue 25094: selecting the bulkhead region

The recorded argument for which region to recover into: capacity for a full Gitaly restore, latency for warm replicas, isolation; settled with a GCP support ticket and a working group, landing on us-central1.

Carry forwardDestination choice is a capacity negotiation with the provider, not a map exercise; open the ticket before the incident.
gitlab.com · production-engineering#25094
SourceGitLabchecked 2026-10

change_regional_recovery issue template

The pre-written C1 change issue an engineer files to begin a regional recovery: services-by-region inventory, data-recovery method per service, and a pre-printed downtime component of 96 hours.

Carry forwardWrite the evacuation declaration before the disaster, as a form; what you cannot pre-fill is what you have not designed.
gitlab.com · production repo issue template
SourceGitLabchecked 2026-10

Zonal and Regional Recovery runbook

The live operator guide: recoveries start with a change issue (on the separate ops instance if GitLab.com is down), HAProxy drains via Consul KV, replica rebuild timings, and an explicit caution that falling back to the recovered zone can cost more downtime.

Carry forwardThe runbook's first step is proof the ops plane is separate; if your first step runs on the thing being recovered, start there.
gitlab.com · runbooks/disaster-recovery/recovery.md
Case studyGitLab2024 to 2026

Recovery measurements: the drill ledger

Timestamped results of repeated gamedays: production Gitaly restore drill at 2 h 05 m with 45 VMs provisioned in 39 minutes (2024-10-21); staging drills from 1 h 15 m to 4 h 15 m, with the slow runs annotated (snapshot quota, boot debugging).

Carry forwardKeep the drill times in a table in the repo, with the embarrassing runs annotated; the variance between drills is your real planning number.
gitlab.com · recovery-measurements.md
Case studyGitLabchecked 2026-10

Gameday runbook and confidence ladder

Near-weekly mock DR events with named change technician and reviewer roles, and a per-service confidence grading in which regional "high confidence" requires infrastructure already standing and ready to receive traffic.

Carry forwardGrade each service's recovery confidence on the four-level ladder; "we have a plan" is the second rung of four, not the top.
gitlab.com · gameday.md
Case studyGitLab2024-10

Change issue 18645: a production gameday, honestly scoped

The drill that produced the production numbers: restoring Gitaly VMs from snapshots in production, with the limit stated in the issue itself: "this is only a test of restoring. No gitaly VMs will be replaced."

Carry forwardA restore drill and a cutover drill are different products; record which one your numbers come from, as this issue does.
gitlab.com · production#18645
SourceGitLab2023-10 to 2025-05

MR 6409: the first gameday plan, closed unmerged

The draft runbook for "the very first gameday FY24Q3" sat unmerged for nineteen months and was closed in May 2025, while the practice it proposed went ahead and evolved past it. The artefact of a plan outlived by its own execution.

Carry forwardDrill cadence beats drill documentation; a rotting draft is evidence the practice is alive somewhere else, and the somewhere else is what to read.
gitlab.com · runbooks!6409
Eng blogDatadog2026-06

When failover isn't safe

A zonal gameday left PostgreSQL clusters with no promotable standby: every candidate exceeded the lag threshold and Patroni "correctly rejected the failover attempt". The fix, synchronous replication for failover candidates only, is benchmarked at +32% to +53% average write latency by commit mode, with DRBD evaluated and rejected.

Carry forwardPrice the write tax per candidate set, not per cluster: sync for the failover candidates, async for read replicas contains the cost.
datadoghq.com · postgresql-ha-kubernetes
PostmortemDatadog2023-05

2023-03-08: one update, five regions, three clouds

An unattended systemd upgrade fired fleet-wide between 06:00 and 07:00 UTC and deleted container routes everywhere at once. The postmortem's own framing: why did one update apply in five regions spanning three cloud providers? Recovery then tripped regional rate limits in the clouds it leaned on.

Carry forwardList everything in your estate that changes on a synchronised clock; each item is a region-spanning failure domain your evacuation cannot escape.
datadoghq.com · multiregion incident
PostmortemDatadog2023-05

2023-03-08: the response deep dive

Internal monitoring alerted three minutes after the first faulty upgrade; a high-severity incident was declared 18 minutes in; the response ran an emergency operations center with roughly 70 incident commanders and hundreds of responders over two days, with full resolution after historical backfill.

Carry forwardFast detection does not shorten a global event much; staffing and recovery ordering dominate once every region needs the same hands.
datadoghq.com · incident-response deep dive
SourcePatroni projectchecked 2026-10

replication_modes.rst

The documentation that states the async bargain exactly: the cluster "is allowed to lose some committed transactions to ensure availability", worst-case bounded by maximum_lag_on_failover bytes plus a recent-writes window; sync and quorum modes remove the loss at a price the Datadog post measures.

Carry forwardRead your HA tool's loss-budget parameter and write its current value into the DR document; that number, not the marketing page, is your RPO.
github.com · patroni/docs/replication_modes.rst
SourceOrchestrator (openark)checked 2026-10

topology-recovery.md

The failover engine GitHub ran in 2018 documents its own restraints: an anti-flapping block period per cluster, recoveries held until a human acknowledges, and a separate force-master-failover path that deliberately ignores the blocks.

Carry forwardGood failover automation ships with a refusal system; configure the block period and the acknowledgement path before the first incident, because the defaults encode someone else's blast radius.
github.com · orchestrator/docs/topology-recovery.md
VendorCockroachDBv25.2, checked 2026-10

Multi-region survival goals

The documented price of never failing over: SURVIVE REGION FAILURE needs at least three database regions, raises replication factor from 3 to 5, and adds at least one inter-region round trip to every write; reads are unaffected.

Carry forwardUse this as the honest lower bound when someone proposes "the database should just survive a region": three regions and a write-latency floor is the entry fee, from the vendor itself.
github.com · cockroachdb/docs survival goals
PostmortemGitHub2018-10 †

October 21 post-incident analysis

43 seconds of partition; Orchestrator promoted West Coast primaries; the healed East held writes the West never saw; 24 hours 11 minutes of degradation while integrity was put ahead of usability. Remediations include keeping Orchestrator from promoting across region boundaries.

Carry forwardThe detector's confidence window (43 s) must be longer than the shortest partition your network can heal, or your failover will race your network team.
github.blog · oct21-post-incident-analysis
PostmortemCloudflare2023-11 †

Control plane and analytics outage post mortem

Total power loss at the primary control-plane facility. Core services reached the DR site in about six hours; "a handful of products did not properly get stood up on our disaster recovery sites", the newer ones without an implemented, tested procedure; log processing waited ~40 hours for the facility itself.

Carry forwardEnrollment in DR must be a launch gate, not a backlog item, because the services that miss it are always the newest ones nobody has drilled.
blog.cloudflare.com · post-mortem
PostmortemAWS2025-10 †

DynamoDB service disruption in US-EAST-1

A race between redundant DNS Enactors emptied the regional endpoint record; the cascade reached EC2's instance-launch workflows, NLB health checks and Lambda across roughly 15 hours, and showed how much global control-plane surface lives in one region.

Carry forwardWalk your evacuation plan and mark every step that calls a control plane homed in the region being evacuated; those steps need a pre-provisioned alternative (static stability) or they are waits.
aws.amazon.com · message/101925
Case studydanluu/post-mortemschecked 2026-10

The curated postmortem ledger

The long-running community index of public postmortems, with one-paragraph mechanism summaries; used here to corroborate the AWS 2025 mechanism (the Enactor race, the EC2 "congestive collapse") and as the fastest index into region-scale incident history.

Carry forwardSearch this index by mechanism, not company, when you need precedent for a design review.
github.com · danluu/post-mortems
PostmortemRoblox2022-01 †

Return to Service 10/28 to 10/31 2021

73 hours down. A Consul streaming feature and a BoltDB issue degraded the single cluster carrying every workload; monitoring depended on the same cluster, so diagnosis ran blind; remediations include multiple datacenters and breaking the observability circularity.

Carry forwardFind your singleton coordination layer (service discovery, feature flags, config store); its blast radius is your real region map.
about.roblox.com · return-to-service
Eng blogNetflix2016-03 †

Global Cloud: Active-Active and Beyond

The account of making any member servable from any region, motivated by removing a single-region point of failure, including a real event where traffic stayed failed over for more than 24 hours and was returned gradually to let services scale and caches warm.

Carry forwardPlan the return leg: failback at cache-cold speed is its own incident, and the professionals do it in hours on purpose.
netflixtechblog.com · active-active and beyond
Eng blogNetflix2013-12 †

Active-Active for Multi-Regional Resiliency

The original decision record in blog form: the Christmas Eve 2012 regional ELB outage bought the multi-region programme, with geo-DNS override tooling to send all users to a healthy region.

Carry forwardPosture changes are bought by incidents; the cheap version is to price the change against someone else's incident instead of waiting for your own.
techblog.netflix.com · active-active, 2013
PaperMeta / USENIX2018-10 †

Maelstrom (OSDI '18)

The drain system: runbooks of traffic shifts that respect inter-system dependencies, draining most web services from a datacenter in under ten minutes, exercised in regular drain tests on production, with 100+ real mitigations over four-plus years. One reported fiber cut: most user-facing traffic out in ~17 minutes, everything in ~1.5 hours.

Carry forwardThe emergency operation and the drill must be the same code path and the same runbook; separate "DR tooling" is tooling that has never run.
usenix.org · osdi18 maelstrom
PaperMeta / SOSP2019-10 †

Taiji (SOSP '19)

Edge-to-datacenter routing as a continuously re-solved assignment problem: fractions of traffic per edge node per datacenter, adjusted for load and failure events; connection-aware grouping of users cut backend query load by 17%.

Carry forwardWhen steering is a table of fractions recomputed continuously, an evacuation is an input change, not a mode change; that is the end state the drill cadence is building toward.
research.facebook.com · taiji
TalkNetflix / QCon2016-06 †

Chaos Kong: endowing Netflix with antifragility

The talk record of failing an entire region's traffic over to another region as a regular production exercise, coordinated by tooling ("Flow") that also handles real recoveries; the drill and the emergency are the same muscle.

Carry forwardA regional failover you have not run this quarter is a hypothesis, and this talk is the citation for making the exercise a calendar item.
qconnewyork.com · chaos kong, 2016
TalkMeta / @Scale2023 †

Warehouse disaster recovery and region shutdown tests

Meta's stated strategy for single-region failure is to drain the region, with an explicit DR capacity buffer in healthy regions; an early region-shutdown test failed because the orchestrator shut down before the data-plane services it managed, which is the kind of ordering bug only a real shutdown finds.

Carry forwardBudget the buffer: an evacuation plan without pre-bought destination capacity is a plan to queue; and expect the first full test to fail on ordering.
atscaleconference.com · warehouse DR
07

Build a miniature, then productionise it

Six rungs. The line from toy to real is crossed at rung four, where the drill starts producing numbers someone could be held to.

Stand up a cluster whose failover can refuse

Run a three-node PostgreSQL cluster under Patroni in containers, two "zones" on one machine. Kill the primary and watch a promotion; then set maximum_lag_on_failover low, hold back a replica, and kill the primary again so Patroni refuses.

Done when: you can produce both outcomes on demand and explain the refusal from the logs.  Teaches: the loss budget is a real parameter with a real guard, not a policy sentence.

Measure your loss budget under induced lag

Drive steady writes, inject latency between primary and replicas (tc netem), and chart replication lag against the failover threshold. Find the write rate at which a zone fault turns into the Datadog stuck state.

Done when: a chart shows lag crossing the budget and the stuck window's duration at two write rates.  Teaches: the same fault that kills the primary disqualifies the candidates; lag is load-dependent, so your RPO is too.

Add a steering layer and drain on command

Put HAProxy or DNS-weighted routing in front of two stacks ("regions"). Implement drain as a config fraction, 100/0 to 0/100, and measure how long connections take to move at several TTL and keepalive settings.

Done when: a timed drain moves all traffic with zero failed requests, and you know which setting dominates the time.  Teaches: steering speed is set by client behaviour you configured months earlier, not by the switch.

Write the change template, then run the gameday against the clock

Copy GitLab's structure: a pre-filled change issue with roles, per-service recovery method, and an honest downtime number. Run the drill on staging end to end, including creating the MRs, and record timings the way recovery-measurements.md does.

Done when: a second engineer, not the author, completes the drill from the template alone, and the measured time is in the repo.  Teaches: the gap between "documented" and "executable by the person on call".

Break the ops plane on purpose

Re-run rung four with the primary "region" denied: no access to its CI, its secrets, its dashboards, its bastion. Every step that fails is a circular dependency of the Roblox or Cloudflare kind. Relocate or pre-stage each one.

Done when: the full drill completes with the primary region's tooling unreachable.  Teaches: which parts of your recovery live inside the thing being recovered; there are always more than the diagram shows.

Do the capacity and control-plane math, then drill in production

Compute the surviving-region load at evacuation (steady load times the shifted fraction, plus reconnection surge), check it against real quotas and rate limits in the destination, and get the buffer approved as a line item. Then take rung four to production on a quiet service, the way GitLab's 2024-10-21 drill did, and put the next one on the calendar.

Done when: a production drain or restore has run this quarter with its timing recorded, and the destination quota request references your math.  Teaches: the difference between a posture you claim and one you hold; also what Meta's buffer and Datadog's thundering herd mean at your scale.

08

Keep hunting

The queries that found this material, grouped by what they surface. The GitLab ones work because the company operates in public; start there for mechanism, then take the vocabulary to the postmortems.

Operators who publish their runbooks

  • site:gitlab.com runbooks disaster-recovery recovery measurements
  • gitlab "gameday" change issue RTO RPO site:gitlab.com
  • "change_regional_recovery" OR "Downtime Component" template
  • "disaster recovery" blueprint "RTO" "RPO" "not yet been validated"

The machinery's own vocabulary

  • "maximum_lag_on_failover" failover "data loss"
  • "anti-flapping" OR "RecoveryPeriodBlockSeconds" orchestrator failover
  • "drain the region" OR "region evacuation" OR "drain tests"
  • "survive region failure" write latency replication factor

Failovers meeting reality

  • "post-incident analysis" failover "promoted" region OR cross-country
  • (postmortem OR "post mortem") "disaster recovery" "had not" tested
  • "thundering herd" recovery "rate limits" postmortem
  • "single Consul cluster" OR "circular dependency" telemetry outage

Paper-to-production chains

  • Maelstrom OSDI 2018 drain interdependent traffic
  • Taiji SOSP 2019 edge traffic assignment "connection-aware"
  • "chaos kong" regional failover exercise production
  • atscaleconference region shutdown test "DR buffer"
09

References

  1. GitLab, Disaster Recovery blueprint (index) GitLab repository, created 2024-01-29, fetched at tag v17.0.0-ee. Checked 2026-10-11.
  2. GitLab, Disaster Recovery blueprint (zonal) GitLab repository, created 2024-01-29. Checked 2026-10-11.
  3. GitLab, Disaster Recovery blueprint (regional): the bulkhead proposal GitLab repository, created 2024-01-29. Checked 2026-10-11.
  4. GitLab, Regional recovery constraints (same document, constraints list) GitLab repository, 2024. Checked 2026-10-11.
  5. GitLab, Issue 25094: Select a region for the regional recovery "bulkhead" gitlab-com/gl-infra/production-engineering, 2024. Checked 2026-10-11.
  6. GitLab, Zonal and Regional Recovery Guide gitlab-com/runbooks, current master. Checked 2026-10-11.
  7. GitLab, change_regional_recovery issue template gitlab-com/gl-infra/production, current master. Checked 2026-10-11.
  8. GitLab, Measuring Recovery Activities (gameday timing ledger) gitlab-com/runbooks, rows 2024-06 to 2026-07. Checked 2026-10-11.
  9. GitLab, Gamedays runbook and confidence levels gitlab-com/runbooks, current master. Checked 2026-10-11.
  10. Datadog, 2023-03-08 Incident: Infrastructure connectivity issue affecting multiple regions Datadog blog, 2023-05-16. Checked 2026-10-11.
  11. AWS, Summary of the Amazon DynamoDB Service Disruption in US-EAST-1 AWS, October 2025. Link verified via retrieval 2026-10-11; host not directly reachable from the research environment.
  12. Datadog, When failover isn't safe: Building high-availability PostgreSQL on Kubernetes Datadog engineering blog, 2026-06-04. Checked 2026-10-11.
  13. GitHub, October 21 post-incident analysis GitHub blog, 2018-10-30. Content via retrieval 2026-10-11; host not directly reachable from the research environment.
  14. Cloudflare, Post Mortem on Cloudflare Control Plane and Analytics Outage Cloudflare blog, 2023-11-04. Content via retrieval 2026-10-11.
  15. Patroni, Replication modes documentation patroni/patroni repository, current master. Checked 2026-10-11.
  16. Orchestrator, Topology recovery documentation openark/orchestrator repository, current master. Checked 2026-10-11.
  17. CockroachDB, Multi-Region Survival Goals cockroachdb/docs repository, v25.2. Checked 2026-10-11.
  18. Roblox, Return to Service 10/28 to 10/31 2021 Roblox newsroom, 2022-01-20. Content via retrieval 2026-10-11.
  19. Netflix, Global Cloud: Active-Active and Beyond Netflix tech blog, 2016-03-30. Content via retrieval 2026-10-11.
  20. Netflix, Active-Active for Multi-Regional Resiliency Netflix tech blog, December 2013. Content via retrieval 2026-10-11.
  21. Veeraraghavan et al., Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently OSDI '18, USENIX, October 2018. Content via retrieval 2026-10-11.
  22. Xu et al., Taiji: Managing Global User Traffic for Large-Scale Internet Services at the Edge SOSP '19, October 2019. Content via retrieval 2026-10-11.
  23. Meta @Scale, Warehouse Disaster Recovery: Batch-Workload Recovery at Scale At Scale conference, December 2023. Content via retrieval 2026-10-11.
  24. Netflix at QCon New York, Chaos Kong: Endowing Netflix with Antifragility QCon New York, June 2016. Content via retrieval 2026-10-11.
  25. Datadog, 2023-03-08 Incident: A deep dive into our incident response Datadog engineering blog, 2023-05-16. Checked 2026-10-11.
  26. GitLab, Change issue 18645: [GPRD] Gitaly DR Gameday Zonal Restore gitlab-com/gl-infra/production, 2024-10-01. Checked 2026-10-11.
  27. GitLab, Change issue 19428: FY26 Q1 HAProxy/Traffic Routing DR Gameday gitlab-com/gl-infra/production, 2025-03-06. Checked 2026-10-11.
  28. GitLab, MR 6409: Draft: gitaly server healthcheck gameday (closed unmerged) gitlab-com/runbooks, opened 2023-10-03, closed 2025-05-12. Checked 2026-10-11.
  29. danluu, post-mortems: a collection of postmortems GitHub repository, rolling. Checked 2026-10-11.