Graceful Degradation in Practice
Deliberately reducing functionality to preserve the core when a dependency fails — a product decision expressed in architecture.
Definition
Rather than failing entirely when a component fails, the system continues with reduced capability. The essential design work is deciding what is essential, and that is a product question with a technical implementation, not the reverse.
The hierarchy of degraded responses
For any component, from best to worst:
- Fresh data — normal operation.
- Cached or stale data, ideally labelled.
- A sensible default — editorially curated content instead of personalised, a standard price instead of a dynamic one.
- Omit the component entirely — render the page without the recommendation strip.
- A clear error for that component, with the rest of the page working.
- Fail the whole request.
Most systems implement only 1 and 6. The value is in the middle, and it must be built deliberately — it does not emerge.
Industry example
Netflix's approach is the reference case, and its instructive feature is that fallbacks are defined per component with an explicit product decision. If personalised recommendations are unavailable, the interface shows popular or curated content rather than an error. The user may not notice; they certainly still get a working product.
The organisational half matters as much as the technical: the fallback behaviour was agreed with product, not invented by engineers at 3am. Someone decided what a degraded experience should look like and it was designed, which is why it looks deliberate rather than broken.
The other half is that degradation paths must be exercised. A fallback that has never run in production has an even chance of being broken, because nothing tests it. This is a large part of why deliberate failure injection exists.
Where degradation is not acceptable
Some operations must fail rather than degrade, and naming them is as important as naming the fallbacks:
- Authorisation. Never fail open on a permission check.
- Payment capture. Better to decline than to guess.
- Anything with a legal or safety consequence.
- Writes that must be durable. Accepting a write you cannot persist is worse than rejecting it.
The rule: degrade reads freely, degrade writes cautiously, and never degrade correctness for money or access.
Failure scenarios
- Fallbacks that call the failing dependency, so the fallback fails too.
- Silent degradation — the system serves stale data indefinitely and nobody is alerted, so a four-day-old price is served for four days.
- Degraded mode more expensive than normal, so the fallback amplifies the incident.
- Fallback data itself stale or wrong, because nothing maintains it.
- No way to tell from the outside that the system is degraded, so incident response starts late.
Interview question
"Your recommendation service is down. Walk me through every level of degraded response and tell me who decides which one you use."