Failure Modes
Slow rather than down, partial, grey, and failing while reporting success.
6 to work through
-
advanced
A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism.
2 min answer -
advanced
A ride-hailing platform has a severe driver shortage in one region. Demand estimation, pricing, dispatch, queueing and driver incentives all interact. How do you prevent an unstable feedback loop?
2 min answer -
advanced
One instance in a fleet of 20 is failing 30% of its requests. Health checks pass, aggregate error rate is 1.5%, no alerts fire. How do you detect and handle this?
2 min answer -
advanced
One instance in a fleet of fifty is returning correct responses very slowly. Health checks pass and it stays in rotation. How do you detect and handle this?
2 min answer -
advanced
Roblox's 2021 outage lasted 73 hours after a service-discovery and key-value cluster degraded under contention. Trace the failure chain, and identify the three architectural properties that turned a degradation into a three-day outage.
2 min answer -
advanced
What happens when a globally distributed real-time communication service loses an entire region - and how should DNS, health checks, traffic steering, failover, connection draining and data consistency interact?
3 min answer
9 terms in this topic
AWS S3 2017: The Blast Radius of a Typo
A mistyped command during routine debugging removed far more capacity than intended, and the affected subsystems had not been restarted in years.
conceptByzantine Fault
A failure in which a component behaves arbitrarily or deceptively — returning wrong results rather than stopping — as distinct from simply crashing.
conceptCircular Observability Dependency
When the tooling used to diagnose a failure runs on the infrastructure that is failing - so the outage removes the ability to see the outage, convert…
conceptControl-Interval Matching
Setting a feedback loop's control interval to the delay of the thing it actuates - because a controller that adjusts faster than the system can respo…
conceptFail-Fast vs Fail-Safe
Whether a component should stop immediately on detecting a problem, or continue in a degraded but safe mode — a choice that depends entirely on which…
conceptGray Failure
A component that is degraded rather than down, passing health checks while serving a portion of requests slowly or incorrectly.
conceptGrey Failure
A component that is degraded rather than down — slow, intermittently erroring, or failing for a subset of operations — which defeats health checks bu…
patternHedged Request
Sending the same request to a second replica after a short delay and using whichever responds first, to cut tail latency caused by unlucky slow servers.
conceptMetastable Failure
A failure state that sustains itself after the original trigger is gone, because the system's own recovery behaviour generates the load keeping it down.
Neighbouring topics
Distributed Systems
General material on partial failure, coordination and distributed reasoning.
CAP & PACELC
What you must give up during a partition, and the latency choice the rest of the time.
Consistency Models
Linearizable, sequential, causal, eventual, and the session guarantees between them.
Idempotency
Making an operation safe to repeat, because a client that times out cannot know.
Retries & Backoff
Exponential backoff, jitter, retry budgets, and how retries become the outage.
Timeouts & Deadlines
Per-hop timeouts that do not compose, and the deadline budget that replaces them.
Circuit Breakers
Failing fast on a broken dependency, and what you fail fast to.
Backpressure & Flow Control
Telling callers to slow down instead of buffering into congestion collapse.
Load Shedding
Rejecting some work deliberately so the rest can be served correctly.
Bulkheads & Isolation
Partitioning resources so one dependency cannot starve the others.
Leader Election
Agreeing who is in charge, and fencing the one who no longer is.
Consensus Protocols
Raft, Paxos and quorums — what they guarantee and what they cost.
Distributed Locking
Mutual exclusion across machines, and why it is harder than it looks.
Distributed Transactions
Two-phase commit, its blocking failure mode, and when it is still reasonable.
Sagas & Compensation
Replacing atomicity with semantic undo, and ordering the irreversible steps last.
Service Discovery
Finding a healthy address for something whose instances are ephemeral.
Messaging & Queues
Decoupling producer from consumer, and the semantics that come with it.
Event Streaming
Retained ordered logs, consumer offsets, partitions and replay.
Clocks & Ordering
Why wall clocks lie, and how logical clocks and versions restore order.