A single-region service already runs across three availability zones. After a four-hour incident at a cloud provider in a neighbouring region, the board asks for a second region. The team's own two-year incident log reads: eleven incidents, nine caused by its own deploys or configuration changes, two by the loss of a single zone. What does a second region add that a third zone does not, and what does that log say about the request?
Show the full answer Hide the answer
What a zone buys and what a region buys
Availability zones are separate power, cooling and network failure domains inside one region, typically tens of kilometres apart and connected at single-digit-millisecond round trips. That distance is the whole point: synchronous replication across zones costs roughly one to two milliseconds per commit, so a zone can be lost with zero data loss and no change to the application. This is why multi-zone is the default and nobody argues about it.
A second region addresses exactly one class of failure: the things all three zones share. The regional control plane that launches instances, the regional endpoints for managed services, the regional identity and quota planes, and the provider's regional network. When those degrade, the zones degrade together, and a third zone adds nothing.
The price is physics. Cross-region round trips run about 60 to 80 ms between US coasts and over 150 ms transatlantic, so synchronous cross-region replication prices every commit at that figure. Almost everyone therefore replicates asynchronously and accepts a non-zero recovery point objective - which means the second region is not a copy of the first, it is a second system with its own consistency story, its own deploy pipeline and its own configuration drift. Data transfer between zones runs around $0.01 per GB in each direction at 2026 list prices in the major clouds; between regions it is several times that, and a warm standby pays for standing capacity before it serves one user.
What the incident log says
Nine of eleven incidents were self-inflicted. A second region does not reduce that number and adds a configuration surface that increases it, because every change now has to land correctly in two places and the failure mode "deployed to one region only" is new. The two zone-loss incidents were already survivable with the architecture in place.
The decision rule: buy availability where your own incidents come from. For this team that is deploy safety - canaried releases, staged configuration rollout, a tested rollback - which buys more availability per engineer-month than a second region by a wide margin. Say that with the log in hand rather than arguing about regions in the abstract.
When a second region is genuinely right
When it is a requirement rather than availability arithmetic: a residency obligation that puts data in a specific jurisdiction, a user population whose round-trip time is the product's problem, or a contract or regulator demanding a documented recovery in another region. Those are not negotiable with an incident log. Note that each is a different topology - residency wants separate stacks, latency wants read replicas near users, recovery wants a standby - and picking the wrong one costs a rebuild.
When this is the wrong answer
If the service is the dependency other teams' recovery plans assume, the arithmetic changes: its unavailability is multiplied across every dependent, and a provider-level regional event takes the whole company down at once rather than one product. Shared infrastructure and payment rails are the common examples. Even then the first deliverable is a tested failover, not a second region - a standby nobody has failed into is a hypothesis, and the incident is a poor time to test it.