Leader Election
Agreeing who is in charge, and fencing the one who no longer is.
6 to work through
-
intermediate Multiple choice
Your company runs two data centres and wants automatic database failover between them. What is missing?
2 min answer -
advanced
A leader-based coordination service loses contact with its followers while the leader itself remains healthy and continues serving. What failure modes emerge, and what prevents split-brain?
2 min answer -
advanced
A leader-based database cluster loses network connectivity between the leader and its followers, but the leader is still running and still accepting writes from application servers that can reach it. What happens, and how does correct leader election prevent it?
3 min answer -
advanced
A nightly job occasionally runs twice, producing duplicate charges. The team proposes a distributed lock. What do you say?
2 min answer -
advanced
A team designs active-active across two regions with automatic failover based on health checks. What is wrong?
2 min answer -
advanced
Your cluster fails over spuriously under load, but a real leader failure takes 45 seconds to detect. How do you resolve the tension?
2 min answer
5 terms in this topic
GitHub 2018: 43 Seconds of Partition, 24 Hours of Recovery
A 43-second network partition triggered an automated cross-region database failover, and reconciling the resulting divergence took over 24 hours.
patternLeader Election in Practice
Selecting one node to coordinate, why majority quorum is non-negotiable, and why fencing at the resource is what actually prevents corruption.
practiceProgress-Based Liveness
Detecting a leader or worker that is alive but not advancing, by checking whether work is progressing rather than whether the process responds.
conceptSplit Brain
A partition in which two subsets of a cluster each believe they are authoritative, accepting conflicting writes that cannot afterwards be reconciled.
conceptSplit Vote
An election in which no candidate obtains a majority, so the term ends with no leader and the process must repeat.
Neighbouring topics
Distributed Systems
General material on partial failure, coordination and distributed reasoning.
CAP & PACELC
What you must give up during a partition, and the latency choice the rest of the time.
Consistency Models
Linearizable, sequential, causal, eventual, and the session guarantees between them.
Idempotency
Making an operation safe to repeat, because a client that times out cannot know.
Retries & Backoff
Exponential backoff, jitter, retry budgets, and how retries become the outage.
Timeouts & Deadlines
Per-hop timeouts that do not compose, and the deadline budget that replaces them.
Circuit Breakers
Failing fast on a broken dependency, and what you fail fast to.
Backpressure & Flow Control
Telling callers to slow down instead of buffering into congestion collapse.
Load Shedding
Rejecting some work deliberately so the rest can be served correctly.
Bulkheads & Isolation
Partitioning resources so one dependency cannot starve the others.
Consensus Protocols
Raft, Paxos and quorums — what they guarantee and what they cost.
Distributed Locking
Mutual exclusion across machines, and why it is harder than it looks.
Distributed Transactions
Two-phase commit, its blocking failure mode, and when it is still reasonable.
Sagas & Compensation
Replacing atomicity with semantic undo, and ordering the irreversible steps last.
Service Discovery
Finding a healthy address for something whose instances are ephemeral.
Messaging & Queues
Decoupling producer from consumer, and the semantics that come with it.
Event Streaming
Retained ordered logs, consumer offsets, partitions and replay.
Clocks & Ordering
Why wall clocks lie, and how logical clocks and versions restore order.
Failure Modes
Slow rather than down, partial, grey, and failing while reporting success.