Correlated Failure
also called Common-Mode Failure, Shared Fate
Failures that are not independent, so redundancy multiplies far less than the arithmetic promises - usually because replicas share code, configuration or a deployment.
Redundancy improves availability only when failures are independent. Three replicas at 99.9% give 99.9999% if and only if one failing does not change the probability of another failing. In practice replicas share almost everything that matters:
- The same code, so a bug triggered by a particular input affects all of them simultaneously.
- The same configuration, distributed by a mechanism designed to be fast and global.
- The same dependencies — one database, one identity service, one DNS resolver.
- The same deployment, so a bad version reaches all of them.
- The same capacity assumptions, so losing one may overload the survivors.
- The same physical substrate, if they share a rack, zone or power domain.
Why it matters
It explains the persistent gap between calculated and observed availability. Teams add replicas, compute an impressive number, and continue to have outages — because redundancy protects against uncorrelated hardware and infrastructure failure, which is real, worthwhile, and not the category causing most modern incidents.
Most serious incidents come from change: a deploy, a configuration push, a dependency change, a traffic shift. Instance redundancy provides no protection against any of them.
Implementation patterns
- Staged rollout with blast-radius limits. Change reaches one instance, then one location, then a region, gated on health. This is the primary defence: it converts "correlated" into "sequential".
- Version diversity during rollout, so a bad version is present only in a minority. Redundancy of identical instances provides no diversity, which is the point.
- Rollback faster than rollout, and independent of what broke — a locally cached last-known-good configuration and an automatic revert on health failure.
- Capacity sized for N-1 or N-2, so survivors can absorb the load rather than falling in sequence.
- Cell isolation, so the unit of failure is a defined subset of users.
- Diversity where it is affordable — different availability zones, different providers for the few dependencies where the cost is justified.
Industry example
Global edge and content platforms are the clearest case, because their redundancy is enormous and their outages are nonetheless usually global. A platform with hundreds of locations and thousands of machines still experiences fleet-wide events, and the cause is almost always a change: a configuration push, a rule update, a software release.
That is why mature edge platforms invest far more in the change pipeline than in the replica count — validation before distribution, staged rollout by blast radius, health-gated promotion using real traffic signals rather than "did it apply", and a rollback path that does not depend on the control plane that may have been broken by the change.
The same reasoning explains why cell-based architectures are adopted at large consumer platforms: cells provide the diversity that identical replicas do not, because a change can be applied to one cell and observed before it reaches the rest.
Failure scenarios
- Availability calculated as if independent, producing a number nobody should believe.
- Configuration treated as low-risk because it is "not code", while being the fastest global change mechanism in the system.
- Redundancy sized at full utilisation, so losing one replica cascades into losing the rest.
- All replicas upgraded simultaneously for operational convenience.
- "Multi-region" that shares a control plane, an identity provider or a configuration source, so the regions fail together for a reason unrelated to geography.
Trade-offs
Reducing correlation costs speed and simplicity. Staged rollouts make every change slower. Version diversity means running mixed versions, which complicates compatibility and debugging. Cells multiply operational surface. N-1 capacity means paying for idle headroom.
Those costs are the price of redundancy that actually works. The alternative is a system with excellent redundancy on paper whose failures are all global.
Interview question
"You have three replicas across three availability zones and you calculate six nines. Tell me three realistic ways all three fail within the same minute, and what you would change for each."