beginner 3 min answer Multiple choice

Three copies of a service run in three availability zones behind one load balancer. A regional event takes all three out inside the same minute although no data centre lost power or cooling. Which property of availability zones explains why the physical separation did not help?

availability-zonesblast-radiuscontrol-planestatic-stabilityfundamentals
Pick one
Show the full answer Hide the answer

What is being tested

Whether you know what a zone physically is, and therefore which failures it is and is not a boundary for. Teams buy a third zone and believe they have bought independence. They have bought independence from one specific class of event.

The mechanism

An availability zone is one or more discrete data centres with separate and redundant power, cooling, networking and connectivity. AWS's published fault-isolation guidance, current in 2026, describes zones inside a region as meaningfully separated — up to roughly 60 miles or 100 km — while staying close enough for single-digit-millisecond networking between them. That geometry is a deliberate compromise: far enough that one substation, one flood plain or one fire does not cover two zones, close enough that synchronous replication costs 1 to 2 ms rather than tens.

Everything that is regional by construction still correlates across all three zones. The control plane that launches instances, mutates load-balancer configuration and issues credentials is regional. Regional service endpoints — object storage, the managed queue, the token service — are regional. So is the configuration you push, the DNS record you change, the container image you roll out and the bug inside it. A three-zone deployment is three copies of the same code, reading the same configuration, through the same endpoints. Replication across zones multiplies the copies and does nothing to the shared inputs.

Why the other options fail

"Not far enough apart." This is the right instinct applied to the wrong failure. The ~100 km ceiling is chosen precisely so that correlated physical events do not span zones, and the stem rules those out: nothing lost power or cooling. Pushing zones further apart would also break the thing the distance buys, because synchronous commit stops being affordable past a few milliseconds of round trip.

"Latency let the replicas diverge." Cross-zone round trips are sub-millisecond to low single-digit milliseconds, which is why multi-zone synchronous commit is the default posture for managed databases. Divergence from latency alone is the cross-region problem at 60 to 80 ms, not the cross-zone one.

"Same building." Attractive because people read "one or more data centres" and invert it. Providers document each zone as physically separate from the others; and even if two were co-located, that would be a provider defect rather than a design decision you could have made differently.

When not to buy another zone

A third or fourth zone is worth the cross-zone traffic and the stranded headroom only if the failures you have actually had were zonal: a power event, a cooling event, a zonal network partition. Pull the last three incidents. If they were a configuration push, a bad deploy or a dependency, another zone changes nothing and the money belongs in static stability instead — enough pre-provisioned capacity in the surviving zones that no control-plane call is needed to survive, and a rehearsed way to take a zone out of rotation. That last part became cheap in 2023: Route 53 Application Recovery Controller zonal shift reached general availability in January 2023 and zonal autoshift followed in November 2023, which turned "drain a zone" from a project into an API call.