Distributed Locking
Mutual exclusion across machines, and why it is harder than it looks.
4 to work through
-
advanced
A media platform uses a distributed lock to stop two workers rendering the same expensive export. Workers sometimes crash while holding the lock, and occasionally two workers process the same job anyway. Diagnose both problems and redesign.
2 min answer -
advanced
A team proposes a distributed lock to stop two workers processing the same transaction. Workers sometimes crash or pause while holding the lock. What is the safer design, and when is a lock genuinely necessary?
2 min answer -
advanced Multiple choice
A team uses a distributed lock in a key-value store to ensure only one worker processes a job. Occasionally two workers process the same job. Explain.
2 min answer -
advanced
Three designs need distributed locks: a nightly report, a per-customer state machine, and a global config reload. For each, is a lock the right answer?
2 min answer
6 terms in this topic
Distributed Locks in Practice
Why lease-based locking is unsafe without fencing, when locks are unavailable entirely, and the partitioning alternative that removes the problem.
patternFencing Token
A monotonically increasing number issued with a lock, checked by the resource, so a holder whose lease expired cannot act on stale authority.
conceptLease
A lock with an expiry, granting exclusive rights for a bounded time so that a crashed holder cannot block the system forever.
conceptLock Lease Expiry
The timeout on a distributed lock that prevents a crashed holder deadlocking the system — and the source of the pattern's hardest failure mode.
patternOptimistic Concurrency Control
Allowing concurrent work without locks and detecting conflict at write time by checking that the underlying version has not changed.
protocolRedlock
An algorithm for distributed locking across independent Redis instances, and the subject of a well-known critique about what locks can guarantee at all.
Neighbouring topics
Distributed Systems
General material on partial failure, coordination and distributed reasoning.
CAP & PACELC
What you must give up during a partition, and the latency choice the rest of the time.
Consistency Models
Linearizable, sequential, causal, eventual, and the session guarantees between them.
Idempotency
Making an operation safe to repeat, because a client that times out cannot know.
Retries & Backoff
Exponential backoff, jitter, retry budgets, and how retries become the outage.
Timeouts & Deadlines
Per-hop timeouts that do not compose, and the deadline budget that replaces them.
Circuit Breakers
Failing fast on a broken dependency, and what you fail fast to.
Backpressure & Flow Control
Telling callers to slow down instead of buffering into congestion collapse.
Load Shedding
Rejecting some work deliberately so the rest can be served correctly.
Bulkheads & Isolation
Partitioning resources so one dependency cannot starve the others.
Leader Election
Agreeing who is in charge, and fencing the one who no longer is.
Consensus Protocols
Raft, Paxos and quorums — what they guarantee and what they cost.
Distributed Transactions
Two-phase commit, its blocking failure mode, and when it is still reasonable.
Sagas & Compensation
Replacing atomicity with semantic undo, and ordering the irreversible steps last.
Service Discovery
Finding a healthy address for something whose instances are ephemeral.
Messaging & Queues
Decoupling producer from consumer, and the semantics that come with it.
Event Streaming
Retained ordered logs, consumer offsets, partitions and replay.
Clocks & Ordering
Why wall clocks lie, and how logical clocks and versions restore order.
Failure Modes
Slow rather than down, partial, grey, and failing while reporting success.