Graceful Degradation
Reduced but useful function when a dependency is gone.
6 to work through
-
intermediate Multiple choice
A recommendation service is unavailable during a high-traffic period. Should the application show popular items, cached recommendations, results from an older model, or nothing at all?
2 min answer -
intermediate
A streaming service's homepage assembles rows from a dozen personalisation services. One returns 4-second latency instead of 40 ms. What happens, and what should have been designed?
2 min answer -
intermediate
The recommendation service is down. Walk through every level of degraded response and say who decides which one is used.
2 min answer -
intermediate
What distinguishes a real graceful-degradation capability from an aspiration, and how do you design and validate a degradation ladder?
3 min answer -
intermediate
What happens to your marketplace if the payment provider is unavailable for two hours during peak?
2 min answer -
advanced
A payments platform's fraud service is degraded during a peak commerce event. Should the payment path wait, skip the check, use cached risk scores, fall back to simpler rules, or fail? What should determine the answer?
2 min answer
4 terms in this topic
Fallback Independence
The requirement that a degraded path fail independently of the primary - and be exercised continuously, because fallbacks that are never run do not work.
practiceFeature Criticality Tiering
Classifying product functionality by whether it must work, should work, or can be dropped, so degradation decisions are made in advance by the business.
patternGraceful Degradation in Practice
Deliberately reducing functionality to preserve the core when a dependency fails — a product decision expressed in architecture.
patternValue-Graduated Fallback
A degradation ladder whose depth depends on the value at stake, so a low-value transaction proceeds on weak signals while a high-value one fails clos…
Neighbouring topics
Reliability & Resilience
General material on designing for failure.
SLI, SLO & SLA
The measurement, the internal target, and the contractual commitment.
Error Budgets
Unreliability as a resource that feature velocity spends.
Availability Mathematics
Series dependencies multiplying, and redundancy that is not independent.
RTO & RPO
How long recovery may take and how much data may be lost.
Redundancy
N+1, N+2 and 2N, correlated failure, and headroom for the failure itself.
Failover
Detection latency, promotion, fencing and the cost of failing over wrongly.
Fault Isolation
Cells, zones, tenants and the partitions that bound a failure.
Chaos Engineering
Hypothesis-driven failure injection with a bounded blast radius.
Game Days
Testing the response — runbooks, access, comms — not only the system.
Incident Management
Command roles, severity levels and mitigation before diagnosis.
Postmortems
Finding the systemic conditions, and completing the actions afterwards.
On-Call
Sustainable rotations, actionable pages and handover discipline.
Capacity Planning
What does not autoscale, and the lead-time items that need a date.
DR Testing
Restore drills, timed against the stated RTO, into a clean environment.
Resilience Testing
Exercising retries, breakers and fallbacks that are otherwise never run.
Static Stability
Serving from existing state when the control plane is unavailable.
Degradation Modes
Deciding in advance what is shed first and what is protected.
Reliability Culture
Blamelessness, error budget policy and reliability as a funded property.