Capacity Planning
What does not autoscale, and the lead-time items that need a date.
4 to work through
-
intermediate
A service runs across three availability zones at 70% CPU during peak. The team wants to survive losing one zone at peak with no user impact. Roughly what does that require, and what does the same arithmetic say about running across two zones?
3 min answer -
advanced
A grocery delivery platform experiences a sudden multi-week increase in demand well beyond any forecast. Which capacity constraints bind first, and which cannot be solved with autoscaling?
2 min answer -
advanced
How should a platform with extreme scheduled peaks decide how much capacity headroom to hold, and what evidence should drive it?
2 min answer -
advanced
Your platform runs across three availability zones. How would you determine whether it actually survives losing one?
2 min answer
2 terms in this topic
Capacity Headroom
The deliberate gap between current load and maximum capacity, sized to absorb growth, spikes and the loss of a failure domain.
metricReliability Headroom
The gap between provisioned capacity and the load that would be carried after the largest planned failure, measured under peak conditions.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.