Degradation Modes
Deciding in advance what is shed first and what is protected.
3 to work through
-
advanced
A crypto aggregator's upstream exchange feeds become unreliable during extreme market volatility - the moment users most need them. How should stale data, provider failover, breakers, caching and user-facing warnings interact?
2 min answer -
advanced
An e-commerce platform designs explicit degradation modes for extreme demand events. What must be true of the mode ladder for it to work on the day?
2 min answer -
advanced
Define three degradation modes for a real-time marketplace. What triggers each, what stops working, and how does the system return to normal?
2 min answer
4 terms in this topic
Correlated Degradation
The condition where a dependency's reliability worsens for the same reason that user demand rises, so the system is least capable exactly when it is …
practiceDegradation Ladder
An ordered, pre-agreed list of features that can be disabled under stress, each with a switch that works without deployment, a named owner, a trigger…
conceptDegradation Mode
A defined, intentional reduced state of service that the system enters under specified conditions, with known behaviour and known exit criteria.
conceptDegradation Modes
Explicitly designed operating modes below "fully working" — declared, testable, and switchable rather than emergent.
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Graceful Degradation
Reduced but useful function when a dependency is gone.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.