Evidence ledger 26 sources Checked 11 Oct 2026

Evidence ledger

One row per claim in Evacuating the region: who can actually leave, and what it costs: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

One row per claim. Tiers follow the implementation-archaeology grading (postmortem > source/adr > casestudy > paper/talk > blog > vendor).

Access note (2026-10-11). This session ran inside a cloud environment whose network policy allows gitlab.com, raw.githubusercontent.com, www.datadoghq.com and cloud.google.com but denies most other hosts. Sources on allowed hosts were fetched directly and quotes are copied from the fetched pages. Sources marked [search-retrieved] sit on denied hosts (github.blog, blog.cloudflare.com, aws.amazon.com, about.roblox.com, netflixtechblog.com, usenix.org, research.fb.com, qconnewyork.com, atscaleconference.com): their content was retrieved through the session's server-side search tool, quotes are limited to fragments that retrieval returned verbatim, and claims beyond those fragments are written as attributed paraphrase. The links are real and were returned by retrieval in this session; they could not be opened directly from this container.

# Org Title Tier Published Checked URL Claim I take from it Supporting quote or figure
1 GitLab Disaster Recovery blueprint (index) adr 2024-01-29 2026-10-11 https://gitlab.com/gitlab-org/gitlab/-/blob/v17.0.0-ee/doc/architecture/blueprints/disaster_recovery/index.md GitLab's FY24 regional recovery target was RTO 96 hours / RPO 2 hours; FY25 target 48 hours / 0 "The FY24 targets were: … Regional: 96 hours (RTO), 2 hours (RPO)… The FY25 targets before cell architecture are: … Regional: 48 hours, 0 minutes"
2 GitLab Disaster Recovery blueprint (index) adr 2024-01-29 2026-10-11 https://gitlab.com/gitlab-org/gitlab/-/blob/v17.0.0-ee/doc/architecture/blueprints/disaster_recovery/index.md Regional recovery is a rebuild from backups and had never been validated end to end "Regional recovery requires a complete rebuild of GitLab.com using backups that are stored in multi-region buckets. The recovery has not yet been validated end-to-end, so we don't know how long the RTO is for a regional failure."
3 GitLab Disaster Recovery blueprint (index) adr 2024-01-29 2026-10-11 https://gitlab.com/gitlab-org/gitlab/-/blob/v17.0.0-ee/doc/architecture/blueprints/disaster_recovery/index.md DR scope is deliberately limited to primary services "We have limited the scope of DR to services that support primary services (Web, API, Git, Pages, Sidekiq, CI, and Registry)." DR excludes AI services, observability, billing, search.
4 GitLab Disaster Recovery blueprint (regional) adr 2024-01-29 2026-10-11 https://gitlab.com/gitlab-org/gitlab/-/blob/v17.0.0-ee/doc/architecture/blueprints/disaster_recovery/regional.md The binding constraints on regional RTO are capacity confidence, edge traffic diversion, ops-plane location, quotas "We don't have confidence that Google can provide us with the capacity we need in a new region, specifically the large amount of SSD necessary to restore all of our customer Git data."; "We have not standardized a way to divert traffic at the edge from 1 region to another."; "Operational infrastructure is located in a single region, us-central1."
5 GitLab Disaster Recovery blueprint (regional) adr 2024-01-29 2026-10-11 https://gitlab.com/gitlab-org/gitlab/-/blob/v17.0.0-ee/doc/architecture/blueprints/disaster_recovery/regional.md The proposed fix is a pre-allocated "regional bulkhead" in the recovery region "To give us a head-start on recovery, we propose a 'regional bulkhead' deployment in a new GCP region… Quotas are set and synced so that we can duplicate all of us-east1 in the new region."
6 GitLab Issue: Select a region for the regional recovery "bulkhead" adr 2024-02 2026-10-11 https://gitlab.com/gitlab-com/gl-infra/production-engineering/-/issues/25094 Recovery-region selection criteria were capacity, latency, isolation, settled with a GCP support ticket "We should take the following into account: 1. capacity (ensuring that compute capacity can handle a full Gitaly recovery) 2. latency … 3. isolation. … Open up a support ticket with GCP to discuss recommendations for region selection"
7 GitLab Runbooks: change_regional_recovery template source master 2026-10-11 https://gitlab.com/gitlab-com/gl-infra/production/-/blob/master/.gitlab/issue_templates/change_regional_recovery.md The pre-printed change issue for a regional recovery budgets 96 hours of downtime and is Criticality 1 "# Production Change - Criticality 1 ~C1 … Downtime Component - … 96 hours"
8 GitLab Runbooks: disaster-recovery/recovery.md source master 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/disaster-recovery/recovery.md Every recovery starts with a change issue; if the primary region is down the issue must be created on the separate ops instance "All recoveries start with a change issue using /change declare … If the us-east1 region is unavailable, it will be necessary to create a change issue on the Ops instance"
9 GitLab Runbooks: disaster-recovery/recovery.md source master 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/disaster-recovery/recovery.md Falling back after the outage is itself risky "When a zonal outage ends, exercise caution in falling back on previously down infrasrtucture. Some components (like Gitaly), may require incur more downtimes when falling back to the old zone."
10 GitLab Runbooks: disaster-recovery/recovery.md source master 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/disaster-recovery/recovery.md Measured replica catch-up cost after zonal loss "As of 2022-12-01, it is expected that it will take approximately 2 hours for the new replica to catch up to the primary using a disk snapshot that is 1 hour old."
11 GitLab Runbooks: disaster-recovery/gameday.md casestudy master 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/disaster-recovery/gameday.md DR drills run almost weekly with named roles, and confidence is graded per service "Mock DR events are simulated almost every week by the Ops team"; regional High Confidence requires "We have infrastructure ready to recieve traffic"
12 GitLab Runbooks: disaster-recovery/recovery-measurements.md casestudy master 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/blob/master/docs/disaster-recovery/recovery-measurements.md Measured gameday times: production Gitaly restore drill 2024-10-21 provisioned 45 VMs in 39 min, whole process 2h05; staging zonal drills ran 1h15 to 4h15 "2024-10-21, GPRD, 00:39:00 … the VM provision time is for 45 Production Gitaly VMs"; "2024-10-21, GPRD, 02:05:00"; staging rows 01:15:00–04:15:00
13 GitLab Change issue 18645: [GPRD] Gitaly DR Gameday Zonal Restore casestudy 2024-10-01 2026-10-11 https://gitlab.com/gitlab-com/gl-infra/production/-/issues/18645 Production gamedays exist but restore to a parallel fleet, not a live cutover "This is only a test of restoring. No gitaly VMs will be replaced."
14 GitLab Change issue 19428: FY26 Q1 HAProxy/Traffic Routing DR Gameday casestudy 2025-03-06 2026-10-11 https://gitlab.com/gitlab-com/gl-infra/production/-/issues/19428 Traffic-routing drills measure against RTO/RPO targets set by a DR working group, on staging "moving traffic away from a single zone in gstg to test our disaster recovery capabilities and measure if we are still within our RTO & RPO targets set by the DR working group"
15 GitLab MR 6409: Draft: gitaly server healthcheck gameday (closed unmerged) source 2023-10-03 2026-10-11 https://gitlab.com/gitlab-com/runbooks/-/merge_requests/6409 The plan for the first gameday sat as a draft for 19 months and was closed unmerged "We need to come up with a plan for the very first gameday FY24Q3." Created 2023-10-03, closed 2025-05-12, state: closed, not merged.
16 Datadog When failover isn't safe: Building high-availability PostgreSQL on Kubernetes blog 2026-06-04 2026-10-11 https://www.datadoghq.com/blog/engineering/postgresql-ha-kubernetes/ A zonal gameday left clusters with no safe failover: every candidate exceeded the lag threshold "Because no replica was sufficiently up to date, failover wasn't safe and the clusters were effectively stuck."; "Patroni correctly rejected the failover attempt. The cluster remained without a safe writable primary not due to Patroni, but because there was no safe promotion candidate."
17 Datadog When failover isn't safe blog 2026-06-04 2026-10-11 https://www.datadoghq.com/blog/engineering/postgresql-ha-kubernetes/ The async default trades durability for availability, and the gameday exposed it "the gameday revealed an uncomfortable truth: In the face of certain network failures, our setup prioritized availability over durability in ways that left us with no safe recovery path."
18 Datadog When failover isn't safe blog 2026-06-04 2026-10-11 https://www.datadoghq.com/blog/engineering/postgresql-ha-kubernetes/ Measured price of making failover safe: synchronous commit raises average write latency 32–53% "The percentage increase in average latency for different synchronous_commit modes was: 53% for remote_apply, 46% for on, 38% for remote_write, 32% for local"
19 Datadog When failover isn't safe blog 2026-06-04 2026-10-11 https://www.datadoghq.com/blog/engineering/postgresql-ha-kubernetes/ DRBD was evaluated and rejected; quorum commit not yet adopted Section "Why we didn't choose DRBD"; "We did not test synchronous replication with quorum commit mode as part of this benchmark."
20 Datadog 2023-03-08 incident: Infrastructure connectivity issue affecting multiple regions postmortem 2023-05-16 2026-10-11 https://www.datadoghq.com/blog/2023-03-08-multiregion-infrastructure-connectivity-issue/ One automatic change hit five regions on three clouds inside an hour, so regional evacuation had nowhere to go "Starting on March 8, 2023, at 06:03 UTC, we experienced an outage that affected the US1, EU1, US3, US4, and US5 Datadog regions across all services."; "Why did the update automatically apply in five distinct regions that span dozens of availability zones and run on three different cloud providers?"
21 Datadog 2023-03-08 incident (same postmortem series) postmortem 2023-05-16 2026-10-11 https://www.datadoghq.com/blog/2023-03-08-multiregion-infrastructure-connectivity-issue/ Mass recovery tripped the destination cloud's regional limits "At the scale of tens of thousands of nodes being replaced at once, it also created a thundering herd that tested the cloud provider's regional rate limits in various ways, none of which was obvious ex ante."
22 Datadog 2023-03-08 incident: A deep dive into our incident response postmortem 2023-05-16 2026-10-11 https://www.datadoghq.com/blog/engineering/2023-03-08-deep-dive-into-incident-response/ Detection was fast even though the trigger was global "We were first alerted to the issue by our internal monitoring, three minutes after the trigger of the first faulty upgrade (March 8, 2023, at 06:00 UTC)." Incident fully resolved 2023-03-10 06:25 UTC after backfill.
23 Patroni project docs/replication_modes.rst source master 2026-10-11 https://github.com/patroni/patroni/blob/master/docs/replication_modes.rst (fetched via raw.githubusercontent.com) Async failover has an explicit, configurable data-loss budget "In asynchronous mode the cluster is allowed to lose some committed transactions to ensure availability."; "the amount of lost data on failover is worst case bounded by maximum_lag_on_failover bytes of transaction log plus the amount that is written in the last ttl seconds"
24 Orchestrator (openark) docs/topology-recovery.md source master 2026-10-11 https://github.com/openark/orchestrator/blob/master/docs/topology-recovery.md (fetched via raw.githubusercontent.com) Production failover tooling ships anti-flapping blocks and human acknowledgement paths "orchestrator avoid flapping (cascading failures causing continuous outage and elimination of resources) by introducing a block period"; manual "force-master-failover" ignores blocks.
25 CockroachDB docs: Multi-Region Survival Goals source v25.2 docs, checked at main 2026-10-11 https://github.com/cockroachdb/docs/blob/main/src/current/v25.2/multiregion-survival-goals.md (fetched via raw.githubusercontent.com) Surviving a region without failover costs a write-latency floor and three regions minimum "write latency will be increased by at least as much as the round-trip time to the nearest region"; "At least three database regions are required to survive region failures."; replication factor raised "from 3 (the default) to 5".
26 GitHub October 21 post-incident analysis [search-retrieved] postmortem 2018-10-30 2026-10-11 https://github.blog/news-insights/company-news/oct21-post-incident-analysis/ A 43-second partition triggered automated cross-country promotion; the service degraded 24h11m; remediation included stopping cross-region promotion Retrieval: 43 seconds of lost connectivity; "24 hours and 11 minutes of service degradation"; Orchestrator promoted West Coast primaries; InfoQ summary of the report lists "preventing Orchestrator from promoting database primaries across region boundaries" as a follow-up.
27 Cloudflare Post Mortem on Cloudflare Control Plane and Analytics Outage [search-retrieved] postmortem 2023-11-04 2026-10-11 https://blog.cloudflare.com/post-mortem-on-cloudflare-control-plane-and-analytics-outage/ The DR that existed worked; the services never put on it were the outage Retrieval: "a handful of products did not properly get stood up on our disaster recovery sites. These tended to be newer products where we had not fully implemented and tested a disaster recovery procedure."; log processing "remained unavailable until we were able to restore PDX-04". Outage 2023-11-02 11:43 UTC; most of control plane serving from DR by 17:57 UTC same day.
28 AWS Summary of the Amazon DynamoDB Service Disruption in US-EAST-1 [search-retrieved] postmortem 2025-10 2026-10-11 https://aws.amazon.com/message/101925/ A DNS race emptied the regional endpoint record and the cascade ran ~15 hours through EC2 and NLB AWS describes "three distinct periods of impact" beginning 2025-10-19 11:48 PM PDT; corroborated by the fetched danluu/post-mortems ledger entry: "Two redundant 'DNS Enactor' processes raced… cascaded into EC2 (DWFM 'congestive collapse'), Network Load Balancer health-check flapping, and Lambda for ~15 hours."
29 danluu/post-mortems Curated postmortem ledger (README) casestudy rolling, checked today 2026-10-11 https://github.com/danluu/post-mortems (fetched via raw.githubusercontent.com) Independent curated corroboration of the AWS 2025 mechanism and of global-scope incidents Entry quoted in row 28; also GitHub Feb 2026 availability report entry: hosted runners failed "in every region" for ~5h53m from one policy change.
30 Roblox Return to Service 10/28-10/31 2021 [search-retrieved] postmortem 2022-01-20 2026-10-11 https://about.roblox.com/newsroom/2022/01/roblox-return-to-service-10-28-10-31-2021 73-hour outage; one Consul cluster was the single point of failure and telemetry depended on it Retrieval: outage lasted 73 hours; a "single Consul cluster" supporting multiple workloads exacerbated impact; monitoring that could have identified the cause depended on the affected systems.
31 Netflix Global Cloud - Active-Active and Beyond [search-retrieved] blog 2016-03-30 2026-10-11 https://netflixtechblog.com/global-cloud-active-active-and-beyond-a0fdfa2c3a45 Netflix built the ability to serve any member from any region, and ran a real evacuation that stayed failed-over for more than a day Retrieval: goal "to create a global cloud where we would be able to serve requests from any member in any AWS region"; a real event kept traffic in the west region for over 24 hours, then shifted back gradually "to let services scale up and caches warm."
32 Netflix Active-Active for Multi-Regional Resiliency [search-retrieved] blog 2013-12 2026-10-11 http://techblog.netflix.com/2013/12/active-active-for-multi-regional.html The active-active investment was triggered by the Christmas Eve 2012 regional ELB outage; geo-DNS override tools direct all users to a healthy region Retrieval: the 2012-12-24 regional ELB outage prompted multi-region active-active; "tools to override geo-DNS and direct all users' traffic to a healthy Region."
33 Meta / USENIX Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently (OSDI '18) [search-retrieved] paper 2018-10 2026-10-11 https://www.usenix.org/conference/osdi18/presentation/veeraraghavan Routine drains make evacuation a minutes-scale operation; drains are run as tests on production Retrieval from paper text: "Maelstrom can drain traffic of most of our web services out of a datacenter in less than 10 minutes without any user-visible, service-level impact."; "in production at Facebook for more than four years… used to mitigate and recover from 100+ datacenter outages." Fiber-cut case: most user-facing traffic drained in ~17 minutes, all traffic in ~1.5 hours (Colyer's summary of the paper).
34 Meta / ACM Taiji: Managing Global User Traffic for Large-Scale Internet Services at the Edge (SOSP '19) [search-retrieved] paper 2019-10 2026-10-11 https://research.facebook.com/publications/taiji-managing-global-user-traffic-for-large-scale-internet-services-at-the-edge/ Steering users between regions is a continuous optimisation problem, not an emergency switch Retrieval: Taiji "models edge-to-datacenter routing as an assignment problem" and "continuously adjusts the routing table to accommodate the dynamics of user traffic and failure events"; connection-aware routing cut backend query load 17%.
35 Netflix / QCon Chaos Kong: Endowing Netflix with Antifragility (QCon New York 2016) [search-retrieved] talk 2016-06 2026-10-11 https://qconnewyork.com/ny2016/ny2016/presentation/chaos-kong-endowing-netflix-antifragility.html Region failover is verified by regularly failing an entire region's traffic over in production Retrieval: Chaos Kong exercises "involved failing over a single region into another region" on a regular basis; the talk describes "Flow," which coordinates recovery and enables periodic verification.
36 Meta / @Scale Warehouse Disaster Recovery & region shutdown test talks (At Scale conference) [search-retrieved] talk 2023-09/12 2026-10-11 https://atscaleconference.com/warehouse-disaster-recovery-batch-workload-recovery-at-scale/ Meta's stated single-region failure strategy is to drain the region, with a capacity buffer; an early region-shutdown test failed in the orchestrator Retrieval: "the strategy for tolerating single-region failure is to drain the region"; Meta added a "DR (disaster recovery) buffer so healthy regions can absorb the shifted load"; an early test failed because "our orchestrator shut down prematurely, before all the data plane services could be reaped and shut down."