concept

Redundancy

Multiple instances of a component so that one failing does not fail the system — valuable exactly to the extent the failures are uncorrelated.

redundancyavailabilitycorrelationfailure-domainscost

Definition

Redundancy means having more of something than the workload requires, so a failure is absorbed rather than felt. It comes in degrees: N (no spare), N+1 (one spare), 2N (a full duplicate), and 2N+1.

The property that decides whether it works

Correlation. Three instances behind a load balancer protect against one instance crashing. They protect against nothing if all three are in the same rack, on the same power feed, running the same version of the code with the same bug, sharing the same configuration, or depending on the same downstream.

Real redundancy requires independent failure domains, and the honest question for any design is: what would take out all copies at once? The answers are usually a shared dependency, a shared deployment, or a shared configuration change — and only the first is what people think of.

The one people miss most: a bad deployment is perfectly correlated across every replica. Instance redundancy provides no protection at all, which is why staged rollout and canaries are a resilience mechanism rather than a delivery convenience.

The layers, and what each protects against

Layer Protects against Does not protect against
Multiple instances Instance failure Bad deploy, dependency failure, config change
Multiple availability zones Zone-level power/network failure Region-wide control plane failure
Multiple regions Regional failure Global configuration change, provider-wide identity outage
Multiple providers Provider failure Enormous complexity cost

Each layer roughly doubles cost and adds coordination complexity. The layer worth buying is the one that addresses the failure you have actually assessed as likely and consequential — not the next one up the list.

Failure scenarios

  • Redundancy with a shared dependency. Three application instances, one database.
  • Redundant capacity that cannot absorb the load. Two instances each at 60% utilisation: when one fails, the other needs 120% and falls over. Redundancy must be sized for the surviving capacity to carry the whole load, and this is a common and expensive oversight.
  • Standby never exercised, so it has drifted and does not work when promoted.
  • Correlated deployment, where a bad release hits every replica simultaneously.
  • Redundancy that increases failure probability — a more complex clustered configuration with more ways to go wrong than the single instance it replaced.

Trade-offs

Bought: tolerance of specific, named failures. Sold: cost roughly proportional to the redundancy factor, plus complexity — coordination, consistency, failover logic — which is itself a source of outages. The mature question is not "how much redundancy" but "which specific failure am I buying protection against, and is that the most likely one?"

Interview question

"Two instances each run at 60% CPU. Is that redundant? What is the failure you have not protected against?"