Multi-Region Architecture
also called Regional Architecture Posture, Cross-Region Topology
Running a system in more than one cloud region, which is a ladder of four distinct postures with different costs and guarantees rather than a single property - and the posture a team needs is rarely the one it asks for.
"We need to go multi-region" hides four different projects. A backup copy in a second region, a pilot light, a warm standby and active-active writes differ by orders of magnitude in cost and operational burden, and they deliver recovery points from hours to zero and recovery times from a day to seconds. The argument in most organisations is not about which posture is right; it is that nobody has said which one is being discussed.
The phrase hides a second thing: regional redundancy is rarely limited by the application. It is limited by everything around it — identity, secrets, images, DNS, certificates, service quotas, the deployment pipeline, and whether a human is authorised to pull the trigger. A region is a serving location; a failover is an exercise of your entire control plane under stress.
Why it matters
A four-hour regional impairment spends about five years of a 99.99% annual error budget in one afternoon, so the business case writes itself — and so does the over-build. A warm standby drilled quarterly delivers more availability than an active-active design nobody rehearsed, because the failure mode of the second is a split-brain reconciliation that lasts days.
The cost gradient is steep. A backup copy is a few per cent of the primary's storage cost; a warm standby roughly doubles database cost and adds a scaled-down fleet; active-active doubles compute, doubles the operational surface, and adds a permanent distributed-data problem that shapes every future feature.
Implementation patterns
- Entity homing. Each customer, account or tenant has exactly one region authorised to write it, so a regional loss affects a known subset of entities and conflicting concurrent writes become structurally impossible.
- Asynchronous replication with a stated recovery point, measured in bytes and seconds and alarmed at a fraction of the target, because the realised recovery point is the lag at the moment of failover.
- Global traffic management as its own layer — DNS, anycast or an edge network — so shifting traffic is a weight change rather than a deployment.
- Static regional independence. Each region holds its own images, secrets, certificates and raised quotas, so failover needs no cross-region control-plane call.
- A written failover trigger that pre-authorises an on-call engineer, because the human decision is usually the largest term in the recovery time, and a rehearsed evacuation scheduled as routine.
Industry example
Roblox's published writeup of its October 2021 outage (January 2022) records a platform of more than 18,000 servers and 170,000 containers running in a single data centre, coordinated by a single shared cluster whose degradation produced a 73-hour outage, with monitoring that depended on the systems that were failing. The company said it was moving to multiple availability zones and data centres.
The instructive part is the ordering. Both aggravating factors they named — one shared coordination cluster and an observability dependency — are fixable without changing regions at all, and a team that jumps straight to a second region leaves both in place.
Failure scenarios
- The executable-versus-servable gap. Traffic can be served from region two but the failover cannot be performed, because the pipeline, the secret store or the identity provider lives only in region one.
- Quota refusal during a drill, so recovery time becomes however long a quota increase takes.
- Silent replication lag. A bulk import pushes lag from 200 ms to 6 minutes and the documented recovery point is wrong for the duration.
- Partition with both regions alive, where anything needing global uniqueness breaks quietly and an automated promotion creates two primaries.
- Failback nobody planned, meaning everything written in region two must be reconciled.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Backup copy in a second region | Protection from regional data loss at a few per cent of cost | Recovery time in hours to a day; recovery point up to the snapshot interval |
| Pilot light | Data continuously replicated; modest standing cost | Everything else provisioned at failover time |
| Warm standby | 15-minute recovery times achievable; the path is exercisable | Roughly double database cost; a second fleet to patch and drill |
| Active-active with entity homing | No failover latency for surviving entities | Permanent distributed-data design; double the operational surface; conflict handling forever |
When not to use it
Stay single-region when you have never measured the cost of a regional outage. Model it — hours of downtime times revenue per hour, plus any contractual penalty — against the annual cost of the proposed posture including the hours the drills consume. For many businesses a four-hour regional event every few years is cheaper than a warm standby, and saying so with numbers beats a diagram.
Stay single-region also when the system is not yet multi-zone-correct: a workload that cannot survive the loss of one availability zone will not survive the loss of a region, and the zone work costs a tenth as much. Where the real requirement is data residency rather than resilience, the answer is a second independent deployment serving a different population, with no failover relationship at all.
Interview question
Q: Leadership asks for multi-region after reading about a competitor's regional outage. You have one meeting. What do you establish before agreeing to anything, and what would you propose instead if the numbers do not support it?
What a strong answer covers: establish the recovery point and time the business actually needs, in writing, and the cost of being down for the alternative duration; establish whether the system survives a zone loss today; inventory what exists only in region one; then name the posture rather than saying "multi-region". Propose the cheaper rung with a drill attached when the numbers do not support more. Name the hidden cost: active-active changes every future feature, because from then on every entity needs a home and every write path needs to know it.
Quick check
Quiz: Which single number determines what you lose in a regional failover? Answer: the replication lag at the instant of failover, not the configured recovery-point target — which is why lag is alarmed in bytes and seconds at a fraction of the target.
Flashcard: Why is a regional failover plan often servable but not executable? — Because identity, secrets, images, certificates, the pipeline or the raised quotas exist only in the primary region, so traffic could be served from region two but the switch cannot be performed.