Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
3 to work through
-
intermediate Multiple choice
A service depends on four components, each independently 99.9% available, called in sequence on every request. What is the resulting availability, and what does this imply about architecture?
2 min answer -
intermediate
A stakeholder requires 99.99% availability. Your service depends on five internal services each offering 99.9%. What do you say?
2 min answer -
advanced
A stakeholder wants a video streaming service at 99.99% availability. Walk through whether that is the right target and what it would take.
2 min answer
3 terms in this topic
Availability Arithmetic
Multiplying dependency availabilities in series and combining redundant components in parallel to derive a system's achievable availability.
conceptAvailability Composition
How the availability of a system follows from its dependencies — multiplying for serial dependencies and improving sharply for redundant ones.
conceptNines Table
The mapping between availability percentages and permitted downtime, and the cost curve that comes with it.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.