Zoom's daily meeting participants rose from roughly 10 million in December 2019 to about 300 million by April 2020. Your platform has a plausible 10x surge ahead and a disaster-recovery plan last tested by tabletop review two years ago. Sequence the move from tabletop to a genuinely tested recovery capability, without an outage.
Show the full answer Hide the answer
The sequence
Each step is reversible, and the ordering is the point: every step produces a real finding before the next step raises the stakes.
- Inventory what recovery actually depends on, from the plan. Walk the document and list every system it assumes: DNS, the identity provider, the secrets store, the artefact registry, the deploy pipeline, the runbook wiki. A recovery plan stored in a wiki that runs in the primary region is the classic finding, and it costs an afternoon to discover.
- Restore one backup to a scratch environment and read the data. Not "the backup job reports success" — restore it, start the application, run a query, compare row counts against production. This is where the highest-value failures are found, and GitLab's January 2017 incident is the canonical reminder: several backup mechanisms had silently not been working. Measure the wall-clock restore time, because it is the number your RTO was invented without.
- Measure the real RPO and RTO, and compare them to the stated ones. Write both pairs down. The gap is the honest state of the capability and is the document that funds the rest of the work.
- Fail over a single stateless service in production, lowest-traffic window, with its traffic shifted to the secondary region and shifted back. Low risk, and it exercises the full apparatus: routing, capacity, credentials, observability in the secondary.
- Fail over a stateful service with a read replica, promoting the replica in the secondary and running read traffic against it. Still reversible. This finds the replication lag reality, which is almost never the dashboard figure under load.
- Rehearse the surge and the failover separately before combining them. Load-test the secondary region to the 10x target while it is not serving, because an untested secondary at 10x is a second outage rather than a recovery. Capacity in a DR region is usually provisioned for today's load, not for the surge that caused the disaster.
- A full regional failover with real traffic, announced, staffed, in a low window, with a tested rollback and a defined abort. This is the first step that could cause an incident, and it comes seventh for that reason.
- Then make it routine. Quarterly, scheduled, unremarkable. A capability tested once is a capability that worked once, and the whole point is to make the exercise boring.
Where things diverge, and how you would know
- Replication lag under load, which determines actual RPO and which is measured during the surge rehearsal rather than assumed from a steady-state graph.
- Configuration drift between regions. The secondary has been receiving changes only when someone remembered. Compare it to the primary mechanically, and from then on apply changes to both through the same pipeline or accept that the secondary is a guess.
- Credentials and certificates in the secondary, which expire silently because nothing uses them. A certificate nobody exercises is a certificate that will be expired on the day.
- Quota in the secondary region. Cloud quotas are per-region, and a failover that needs 10x the instances will hit an account limit, which is a support ticket during an outage.
The point of no return
Step 7, and only while traffic is shifted. Everything before it is a measurement. The thing that is genuinely irreversible is data divergence: once both regions accept writes, reconciliation is the hard problem. So the plan must be unambiguous about which region owns writes at every moment, with a mechanism — a lease, a fencing token, a single writable endpoint — rather than a procedure people follow.
How long it really takes
Steps 1–3 are two to three weeks and return most of the value, which is the thing to say to anyone impatient with the sequence. Steps 4–6 are a quarter. Step 7 needs executive sign-off and a staffed window, and a credible plan is two quarters to a tested regional failover, then quarterly thereafter.
When not to do all of this
If the honest requirement is a four-hour RTO and a one-hour RPO, stop after step 3 and automate restore-and-verify. A tested restore path plus backups verified nightly meets that requirement for a fraction of the cost of an active secondary region, and the multi-region work above is justified by an RTO measured in minutes.
Steps 4 to 7 also presuppose a second region that is already provisioned and reasonably current. If the secondary does not exist, building it to pass a test is the wrong order — size it from the measured RTO requirement in step 3, which may conclude that a cold standby restored from verified backups is the correct architecture. Many organisations discover at step 3 that they bought an active secondary they did not need and still cannot restore a database.