Static Stability
Serving from existing state when the control plane is unavailable.
5 to work through
-
intermediate
A service fetches a feature flag value and a database credential from central services on every request. What is wrong?
2 min answer -
intermediate
During a peak event, a platform's configuration service becomes unavailable. What should happen to the serving path, and what design property is being tested?
2 min answer -
intermediate
Explain static stability, and give two examples of designs that violate it.
2 min answer -
advanced Multiple choice
A marketplace's failover mechanism requires calling a control plane to provision replacement capacity. Why is this a design flaw, and what is the alternative?
2 min answer -
advanced
Your disaster recovery plan provisions capacity in the secondary region at failover time. Why is that a problem?
2 min answer
3 terms in this topic
AWS: Static Stability Across Availability Zones
AWS designs services to keep working with the capacity they already have when a zone fails, rather than needing the control plane to provision replac…
patternPre-Provisioned Capacity
Running capacity that is already in place to absorb a failure, rather than depending on a control plane to create it during the failure.
conceptStatic Stability
A system that keeps working correctly in its current state when its control plane or dependencies become unavailable, rather than needing them to sta…
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.