Distributed Systems
General material on partial failure, coordination and distributed reasoning.
6 to work through
-
intermediate
An order service must notify inventory, billing, shipping and analytics when an order is placed. Synchronous calls or events? Justify your choice per consumer.
2 min answer -
intermediate
Design an order submission API that is safe when the client cannot tell whether its request succeeded. What exactly do you store, and when?
2 min answer -
intermediate Multiple choice
You move a user profile service to eventual consistency and support tickets start arriving: users update their name and the old one is still shown. Fix it without abandoning the architecture.
2 min answer -
advanced
A 43-second network partition caused GitHub over 24 hours of degraded service in 2018. How does a 43-second event become a day-long incident?
2 min answer -
advanced Multiple choice
A card payment authorisation service runs active-active across two regions. A network partition splits them. Do you keep accepting authorisations, and what breaks either way?
2 min answer -
advanced
A downstream service slows from 50 ms to 3 s. Within two minutes every service in the request path is down, including ones that do not call it. Explain the mechanism and how you would have prevented it.
2 min answer
24 terms in this topic
Backpressure
A mechanism by which a component under load tells its callers to slow down, rather than accepting work it cannot complete.
patternBulkhead
Partitioning resources so that exhaustion caused by one dependency or tenant cannot starve the others.
conceptCAP Theorem
During a network partition a distributed system must choose between consistency and availability; it cannot have both.
patternCircuit Breaker
A proxy that stops calling a failing dependency after a failure threshold, failing fast instead, and periodically tests whether it has recovered.
conceptConsistent Hashing
A hashing scheme where adding or removing a node remaps only a small fraction of keys, instead of nearly all of them.
toolEvent Stream
An append-only, retained log of events that many independent consumers read at their own position, and can re-read.
conceptEventual Consistency
A guarantee that replicas will converge to the same value if updates stop, with no bound on how long reads may be stale.
patternExponential Backoff
Increasing the wait between retries geometrically, with random jitter, so that failures do not synchronise into a stampede.
conceptFan-Out
One incoming request causing many outgoing ones, which multiplies both load and tail latency.
conceptFault Tolerance
Continuing to operate correctly despite the failure of some components, by design rather than by luck.
practiceGraceful Degradation
Continuing to deliver reduced but useful function when a dependency fails, instead of failing the whole request.
conceptIdempotency
The property that performing an operation many times has the same effect as performing it once.
patternLeader Election
The process by which a group of nodes agrees which one of them is currently in charge of a task that must not run twice.
patternLoad Shedding
Deliberately rejecting a portion of incoming work during overload so that the remainder can be served correctly.
toolMessage Queue
A store that holds messages until a consumer processes them, decoupling producer availability and rate from consumer availability and rate.
conceptQuorum
A minimum number of nodes that must acknowledge an operation for it to count, chosen so that read and write sets are guaranteed to overlap.
patternSaga
A sequence of local transactions across services where each step has a compensating action that semantically undoes it if a later step fails.
conceptScalability
The ability to handle growing load by adding resources, ideally with cost rising no faster than the load.
conceptService Discovery
The mechanism by which a caller finds a currently healthy network address for a service whose instances are ephemeral.
patternShuffle Sharding
Assigning each customer a random combination of workers rather than a fixed shard, so that any two customers rarely share their whole set.
conceptSplit Brain
A partition in which two halves of a cluster each believe they are authoritative, and both accept writes.
conceptThundering Herd
A large number of clients acting simultaneously because they were synchronised by a shared event, producing a spike that the steady-state design neve…
patternTimeout Budget
Assigning a request an overall deadline at the edge and passing the remaining time down each hop, so no service works on something already out of time.
protocolTwo-Phase Commit
A blocking protocol for atomic commit across several resources: a coordinator asks all participants to prepare, then tells them all to commit or abort.
Neighbouring topics
CAP & PACELC
What you must give up during a partition, and the latency choice the rest of the time.
No content yetConsistency Models
Linearizable, sequential, causal, eventual, and the session guarantees between them.
No content yetIdempotency
Making an operation safe to repeat, because a client that times out cannot know.
No content yetRetries & Backoff
Exponential backoff, jitter, retry budgets, and how retries become the outage.
No content yetTimeouts & Deadlines
Per-hop timeouts that do not compose, and the deadline budget that replaces them.
No content yetCircuit Breakers
Failing fast on a broken dependency, and what you fail fast to.
No content yetBackpressure & Flow Control
Telling callers to slow down instead of buffering into congestion collapse.
No content yetLoad Shedding
Rejecting some work deliberately so the rest can be served correctly.
No content yetBulkheads & Isolation
Partitioning resources so one dependency cannot starve the others.
No content yetLeader Election
Agreeing who is in charge, and fencing the one who no longer is.
No content yetConsensus Protocols
Raft, Paxos and quorums — what they guarantee and what they cost.
No content yetDistributed Locking
Mutual exclusion across machines, and why it is harder than it looks.
No content yetDistributed Transactions
Two-phase commit, its blocking failure mode, and when it is still reasonable.
No content yetSagas & Compensation
Replacing atomicity with semantic undo, and ordering the irreversible steps last.
No content yetService Discovery
Finding a healthy address for something whose instances are ephemeral.
No content yetMessaging & Queues
Decoupling producer from consumer, and the semantics that come with it.
No content yetEvent Streaming
Retained ordered logs, consumer offsets, partitions and replay.
No content yetClocks & Ordering
Why wall clocks lie, and how logical clocks and versions restore order.
No content yetFailure Modes
Slow rather than down, partial, grey, and failing while reporting success.
No content yet