Timeouts & Deadlines
Per-hop timeouts that do not compose, and the deadline budget that replaces them.
5 to work through
-
intermediate
A service becomes unresponsive during an incident in a dependency it barely uses. Investigation finds the client had no read timeout. Explain the full mechanism.
2 min answer -
intermediate
Every service in your estate uses the client library's default timeout of 30 seconds. Why is that dangerous, and what would you set instead?
2 min answer -
advanced
A financial-data aggregator depends on hundreds of external banking APIs, some slow, some intermittently down, all with different behaviour. How should timeouts, retries, circuit breakers, caching, stale data and provider isolation interact?
2 min answer -
advanced Multiple choice
A travel search fans out to 40 suppliers with latencies from 80 ms to 4 seconds and varying reliability. Which timeout strategy should the aggregator use?
2 min answer -
advanced
Design the timeout configuration for a request that passes through gateway, orders, pricing and inventory. What numbers, and what rule generates them?
2 min answer
6 terms in this topic
Ambiguous Outcome
The state a system is in when a call times out - not failure, but unknown - and the design obligation to have somewhere to put it.
conceptCascading Timeout
The effect of independently-chosen per-hop timeouts summing to a total far longer than any caller is willing to wait.
conceptDeadline Exceeded
The error returned when a request's overall budget expires — semantically distinct from a per-hop timeout, and a signal that must not be retried blindly.
patternDeadline Propagation
Carrying the originating request's absolute deadline through every downstream hop so that each service works with the remaining budget - and refuses …
patternDeadline Propagation
Passing the remaining time budget down each call in a request chain so downstream services never work on a request whose caller has already given up.
conceptRead Timeout
The bound on how long a client waits for response data after a connection is established — distinct from the connect timeout, and the one that usuall…
Neighbouring topics
Distributed Systems
General material on partial failure, coordination and distributed reasoning.
CAP & PACELC
What you must give up during a partition, and the latency choice the rest of the time.
Consistency Models
Linearizable, sequential, causal, eventual, and the session guarantees between them.
Idempotency
Making an operation safe to repeat, because a client that times out cannot know.
Retries & Backoff
Exponential backoff, jitter, retry budgets, and how retries become the outage.
Circuit Breakers
Failing fast on a broken dependency, and what you fail fast to.
Backpressure & Flow Control
Telling callers to slow down instead of buffering into congestion collapse.
Load Shedding
Rejecting some work deliberately so the rest can be served correctly.
Bulkheads & Isolation
Partitioning resources so one dependency cannot starve the others.
Leader Election
Agreeing who is in charge, and fencing the one who no longer is.
Consensus Protocols
Raft, Paxos and quorums — what they guarantee and what they cost.
Distributed Locking
Mutual exclusion across machines, and why it is harder than it looks.
Distributed Transactions
Two-phase commit, its blocking failure mode, and when it is still reasonable.
Sagas & Compensation
Replacing atomicity with semantic undo, and ordering the irreversible steps last.
Service Discovery
Finding a healthy address for something whose instances are ephemeral.
Messaging & Queues
Decoupling producer from consumer, and the semantics that come with it.
Event Streaming
Retained ordered logs, consumer offsets, partitions and replay.
Clocks & Ordering
Why wall clocks lie, and how logical clocks and versions restore order.
Failure Modes
Slow rather than down, partial, grey, and failing while reporting success.