Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
3 to work through
-
intermediate Multiple choice
Two service instances each run at 60% CPU behind a load balancer. Is that redundant?
2 min answer -
advanced
A platform runs three redundant instances of a component and calculates its availability as extremely high. In practice, all three fail together during incidents. What is the flaw in the reasoning?
2 min answer -
advanced
A platform runs three replicas of every service across three availability zones and still experiences total outages. What kinds of failure does that redundancy not address?
2 min answer
4 terms in this topic
Correlated Failure
Failures that are not independent, so redundancy multiplies far less than the arithmetic promises - usually because replicas share code, configuratio…
patternN+1 and 2N Redundancy
Provisioning one spare beyond required capacity versus provisioning double, and the failure assumptions each encodes.
conceptRedundancy
Multiple instances of a component so that one failing does not fail the system — valuable exactly to the extent the failures are uncorrelated.
practiceShared-Cause Analysis
Auditing a redundant design by asking what every replica has in common, because redundancy protects against independent failure and not against anyth…
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.