Evacuating the region: who can actually leave, and what it costs
How production systems actually move off a failing region or datacenter: the steering, the data-loss budget, the decision authority, and why the rehearsed recovery is measured in days for some organisations and minutes for others.
Reconstructs regional evacuation from GitLab's public DR blueprints, runbooks and gameday timing ledger, Meta's Maelstrom and Taiji papers, Netflix's active-active record, Datadog's failover-safety work, and the postmortems of GitHub 2018, Cloudflare 2023, Roblox 2021, Datadog 2023 and AWS 2025. After reading, an architect can name their organisation's posture (rebuilder, failover-ready, traffic shifter), price the move to the next posture, and audit the two components whose absence silently makes any system a rebuilder: a steering layer outside the region and an ops plane that shares no fate with it.
GitLab's regional-recovery change template pre-prints a 96-hour downtime component, and the biggest failover disaster in the record (GitHub 2018) came from the machinery running exactly as designed; the guard that refuses to fail over (Patroni's lag gate, Orchestrator's anti-flapping block) is the most load-bearing feature in the corpus.
What you get out of it
- The public record splits organisations into rebuilders (days), failover-ready (the disaster zone), and traffic shifters (minutes); the time to evacuate is a property of the posture, not the incident.
- Automated failover belongs inside the boundary you drill at that cadence; GitHub 2018's remediation was to stop Orchestrator promoting across regions, while Meta automates drains precisely because the same code path runs weekly as a test.
- An async replica is not a failover target during the failures that matter: the same fault that kills the primary inflates the lag that disqualifies the candidates (Datadog's stuck gameday), and fixing it cost a measured 46% in write latency.
- Regions partition the request path, not the change path or the control plane: Datadog's one update hit five regions on three clouds in an hour, and AWS's us-east-1 event stalled nominally regional recoveries worldwide.
- No postmortem in this corpus describes a deliberate regional evacuation that ran and failed; the record cannot prove your evacuation works, only your own production drill can.
Scope
Why this, now. The October 2025 us-east-1 cascade and Cloudflare's 2023 control-plane outage restarted the multi-region argument in most design reviews, while GitLab's gameday ledger and Datadog's 2026 failover-safety post have quietly published the real numbers both sides of that argument need.
What it does not cover. Multi-writer database internals (Spanner/CockroachDB get a price tag, not a chapter), DNS and anycast mechanics, compliance-driven DR paperwork, and restore-from-corruption, which is a backup problem covered by an earlier guide.
Other field guides
Coming back from cold: the restart is a harder load case than the crash
Reconstructs, from the Kubernetes project's own design records, the fixes HashiCorp shipped after October 2021, the Chubby and Physalia papers, Netfl…
20 sources · 10 organisations · 4 postmortemsThe way back is also a deploy
Reconstructs rollback as it actually works in production: a forward deploy of an old artifact, underwritten by a compatibility contract fixed weeks b…
28 sources · 25 organisations · 6 postmortemsWhere resilience policy lives: Netflix, 2016 to 2026
Between 2016 and 2026 almost every Netflix library that decided something about the network was retired, while the libraries that decide something ab…
22 sources · 4 organisations · 4 postmortems