Every source behind this page, graded. Cards marked with a dagger (†)
sit on hosts this research environment could not open directly; their content was read
through search-engine retrieval, quotes are limited to what retrieval returned verbatim,
and the ledger (sources.md, shipped beside this page) records the detail.
ADRGitLab2024-01
Disaster Recovery blueprint
The working design document for GitLab.com DR: scope limited to primary services, FY24
regional targets of 96 h RTO / 2 h RPO moving to 48 h / 0, and the admission that regional
recovery "has not yet been validated end-to-end, so we don't know how long the RTO is".
Carry forwardWrite the targets you can defend, then publish the gap between target and validation; the blueprint's honesty is what makes the rest of it credible.
gitlab.com/gitlab-org/gitlab · blueprints/disaster_recovery
ADRGitLab2024-01
Regional recovery proposal (the bulkhead)
Lists eleven concrete constraints on regional RTO, including no standard way to divert
edge traffic, ops tooling in a single region, and no confidence the cloud can supply
recovery capacity; proposes a pre-allocated "regional bulkhead" with synced quotas.
Carry forwardThe eleven-item list is a ready-made audit: run it against your own estate and the items you cannot answer are your real RTO.
gitlab.com · disaster_recovery/regional.md
ADRGitLab2024-02
Issue 25094: selecting the bulkhead region
The recorded argument for which region to recover into: capacity for a full Gitaly
restore, latency for warm replicas, isolation; settled with a GCP support ticket and a
working group, landing on us-central1.
Carry forwardDestination choice is a capacity negotiation with the provider, not a map exercise; open the ticket before the incident.
gitlab.com · production-engineering#25094
SourceGitLabchecked 2026-10
change_regional_recovery issue template
The pre-written C1 change issue an engineer files to begin a regional recovery:
services-by-region inventory, data-recovery method per service, and a pre-printed
downtime component of 96 hours.
Carry forwardWrite the evacuation declaration before the disaster, as a form; what you cannot pre-fill is what you have not designed.
gitlab.com · production repo issue template
SourceGitLabchecked 2026-10
Zonal and Regional Recovery runbook
The live operator guide: recoveries start with a change issue (on the separate ops
instance if GitLab.com is down), HAProxy drains via Consul KV, replica rebuild timings,
and an explicit caution that falling back to the recovered zone can cost more downtime.
Carry forwardThe runbook's first step is proof the ops plane is separate; if your first step runs on the thing being recovered, start there.
gitlab.com · runbooks/disaster-recovery/recovery.md
Case studyGitLab2024 to 2026
Recovery measurements: the drill ledger
Timestamped results of repeated gamedays: production Gitaly restore drill at 2 h 05 m
with 45 VMs provisioned in 39 minutes (2024-10-21); staging drills from 1 h 15 m to
4 h 15 m, with the slow runs annotated (snapshot quota, boot debugging).
Carry forwardKeep the drill times in a table in the repo, with the embarrassing runs annotated; the variance between drills is your real planning number.
gitlab.com · recovery-measurements.md
Case studyGitLabchecked 2026-10
Gameday runbook and confidence ladder
Near-weekly mock DR events with named change technician and reviewer roles, and a
per-service confidence grading in which regional "high confidence" requires
infrastructure already standing and ready to receive traffic.
Carry forwardGrade each service's recovery confidence on the four-level ladder; "we have a plan" is the second rung of four, not the top.
gitlab.com · gameday.md
Case studyGitLab2024-10
Change issue 18645: a production gameday, honestly scoped
The drill that produced the production numbers: restoring Gitaly VMs from snapshots in
production, with the limit stated in the issue itself: "this is only a test of
restoring. No gitaly VMs will be replaced."
Carry forwardA restore drill and a cutover drill are different products; record which one your numbers come from, as this issue does.
gitlab.com · production#18645
SourceGitLab2023-10 to 2025-05
MR 6409: the first gameday plan, closed unmerged
The draft runbook for "the very first gameday FY24Q3" sat unmerged for nineteen months
and was closed in May 2025, while the practice it proposed went ahead and evolved past
it. The artefact of a plan outlived by its own execution.
Carry forwardDrill cadence beats drill documentation; a rotting draft is evidence the practice is alive somewhere else, and the somewhere else is what to read.
gitlab.com · runbooks!6409
Eng blogDatadog2026-06
When failover isn't safe
A zonal gameday left PostgreSQL clusters with no promotable standby: every candidate
exceeded the lag threshold and Patroni "correctly rejected the failover attempt". The
fix, synchronous replication for failover candidates only, is benchmarked at +32% to
+53% average write latency by commit mode, with DRBD evaluated and rejected.
Carry forwardPrice the write tax per candidate set, not per cluster: sync for the failover candidates, async for read replicas contains the cost.
datadoghq.com · postgresql-ha-kubernetes
PostmortemDatadog2023-05
2023-03-08: one update, five regions, three clouds
An unattended systemd upgrade fired fleet-wide between 06:00 and 07:00 UTC and deleted
container routes everywhere at once. The postmortem's own framing: why did one update
apply in five regions spanning three cloud providers? Recovery then tripped regional
rate limits in the clouds it leaned on.
Carry forwardList everything in your estate that changes on a synchronised clock; each item is a region-spanning failure domain your evacuation cannot escape.
datadoghq.com · multiregion incident
PostmortemDatadog2023-05
2023-03-08: the response deep dive
Internal monitoring alerted three minutes after the first faulty upgrade; a
high-severity incident was declared 18 minutes in; the response ran an emergency
operations center with roughly 70 incident commanders and hundreds of responders over
two days, with full resolution after historical backfill.
Carry forwardFast detection does not shorten a global event much; staffing and recovery ordering dominate once every region needs the same hands.
datadoghq.com · incident-response deep dive
SourcePatroni projectchecked 2026-10
replication_modes.rst
The documentation that states the async bargain exactly: the cluster "is allowed to
lose some committed transactions to ensure availability", worst-case bounded by
maximum_lag_on_failover bytes plus a recent-writes window; sync and quorum modes remove
the loss at a price the Datadog post measures.
Carry forwardRead your HA tool's loss-budget parameter and write its current value into the DR document; that number, not the marketing page, is your RPO.
github.com · patroni/docs/replication_modes.rst
SourceOrchestrator (openark)checked 2026-10
topology-recovery.md
The failover engine GitHub ran in 2018 documents its own restraints: an anti-flapping
block period per cluster, recoveries held until a human acknowledges, and a separate
force-master-failover path that deliberately ignores the blocks.
Carry forwardGood failover automation ships with a refusal system; configure the block period and the acknowledgement path before the first incident, because the defaults encode someone else's blast radius.
github.com · orchestrator/docs/topology-recovery.md
VendorCockroachDBv25.2, checked 2026-10
Multi-region survival goals
The documented price of never failing over: SURVIVE REGION FAILURE needs at least
three database regions, raises replication factor from 3 to 5, and adds at least one
inter-region round trip to every write; reads are unaffected.
Carry forwardUse this as the honest lower bound when someone proposes "the database should just survive a region": three regions and a write-latency floor is the entry fee, from the vendor itself.
github.com · cockroachdb/docs survival goals
PostmortemGitHub2018-10 †
October 21 post-incident analysis
43 seconds of partition; Orchestrator promoted West Coast primaries; the healed East
held writes the West never saw; 24 hours 11 minutes of degradation while integrity was
put ahead of usability. Remediations include keeping Orchestrator from promoting across
region boundaries.
Carry forwardThe detector's confidence window (43 s) must be longer than the shortest partition your network can heal, or your failover will race your network team.
github.blog · oct21-post-incident-analysis
PostmortemCloudflare2023-11 †
Control plane and analytics outage post mortem
Total power loss at the primary control-plane facility. Core services reached the DR
site in about six hours; "a handful of products did not properly get stood up on our
disaster recovery sites", the newer ones without an implemented, tested procedure; log
processing waited ~40 hours for the facility itself.
Carry forwardEnrollment in DR must be a launch gate, not a backlog item, because the services that miss it are always the newest ones nobody has drilled.
blog.cloudflare.com · post-mortem
PostmortemAWS2025-10 †
DynamoDB service disruption in US-EAST-1
A race between redundant DNS Enactors emptied the regional endpoint record; the
cascade reached EC2's instance-launch workflows, NLB health checks and Lambda across
roughly 15 hours, and showed how much global control-plane surface lives in one region.
Carry forwardWalk your evacuation plan and mark every step that calls a control plane homed in the region being evacuated; those steps need a pre-provisioned alternative (static stability) or they are waits.
aws.amazon.com · message/101925
Case studydanluu/post-mortemschecked 2026-10
The curated postmortem ledger
The long-running community index of public postmortems, with one-paragraph mechanism
summaries; used here to corroborate the AWS 2025 mechanism (the Enactor race, the EC2
"congestive collapse") and as the fastest index into region-scale incident history.
Carry forwardSearch this index by mechanism, not company, when you need precedent for a design review.
github.com · danluu/post-mortems
PostmortemRoblox2022-01 †
Return to Service 10/28 to 10/31 2021
73 hours down. A Consul streaming feature and a BoltDB issue degraded the single
cluster carrying every workload; monitoring depended on the same cluster, so diagnosis
ran blind; remediations include multiple datacenters and breaking the observability
circularity.
Carry forwardFind your singleton coordination layer (service discovery, feature flags, config store); its blast radius is your real region map.
about.roblox.com · return-to-service
Eng blogNetflix2016-03 †
Global Cloud: Active-Active and Beyond
The account of making any member servable from any region, motivated by removing a
single-region point of failure, including a real event where traffic stayed failed over
for more than 24 hours and was returned gradually to let services scale and caches warm.
Carry forwardPlan the return leg: failback at cache-cold speed is its own incident, and the professionals do it in hours on purpose.
netflixtechblog.com · active-active and beyond
Eng blogNetflix2013-12 †
Active-Active for Multi-Regional Resiliency
The original decision record in blog form: the Christmas Eve 2012 regional ELB outage
bought the multi-region programme, with geo-DNS override tooling to send all users to a
healthy region.
Carry forwardPosture changes are bought by incidents; the cheap version is to price the change against someone else's incident instead of waiting for your own.
techblog.netflix.com · active-active, 2013
PaperMeta / USENIX2018-10 †
Maelstrom (OSDI '18)
The drain system: runbooks of traffic shifts that respect inter-system dependencies,
draining most web services from a datacenter in under ten minutes, exercised in regular
drain tests on production, with 100+ real mitigations over four-plus years. One reported
fiber cut: most user-facing traffic out in ~17 minutes, everything in ~1.5 hours.
Carry forwardThe emergency operation and the drill must be the same code path and the same runbook; separate "DR tooling" is tooling that has never run.
usenix.org · osdi18 maelstrom
PaperMeta / SOSP2019-10 †
Taiji (SOSP '19)
Edge-to-datacenter routing as a continuously re-solved assignment problem: fractions
of traffic per edge node per datacenter, adjusted for load and failure events;
connection-aware grouping of users cut backend query load by 17%.
Carry forwardWhen steering is a table of fractions recomputed continuously, an evacuation is an input change, not a mode change; that is the end state the drill cadence is building toward.
research.facebook.com · taiji
TalkNetflix / QCon2016-06 †
Chaos Kong: endowing Netflix with antifragility
The talk record of failing an entire region's traffic over to another region as a
regular production exercise, coordinated by tooling ("Flow") that also handles real
recoveries; the drill and the emergency are the same muscle.
Carry forwardA regional failover you have not run this quarter is a hypothesis, and this talk is the citation for making the exercise a calendar item.
qconnewyork.com · chaos kong, 2016
TalkMeta / @Scale2023 †
Warehouse disaster recovery and region shutdown tests
Meta's stated strategy for single-region failure is to drain the region, with an
explicit DR capacity buffer in healthy regions; an early region-shutdown test failed
because the orchestrator shut down before the data-plane services it managed, which is
the kind of ordering bug only a real shutdown finds.
Carry forwardBudget the buffer: an evacuation plan without pre-bought destination capacity is a plan to queue; and expect the first full test to fail on ordering.
atscaleconference.com · warehouse DR