Term Kind Topic What it is
Lease concept Distributed Locking A lock with an expiry, granting exclusive rights for a bounded time so that a crashed holder cannot block the system forever.
Linearizability Atomic Consistency, Strong Consistency concept Consistency Models The strongest single-object guarantee — every operation appears to take effect instantaneously at some point between its call and its return.
LinkedIn: Kafka and the Unified Log The Log, Kafka Origin case-study Event Streaming LinkedIn replaced a tangle of point-to-point data pipelines with a single durable log, turning an O(n²) integration problem into an O(n) one.
Load Shedding pattern Load Shedding Deliberately rejecting a portion of incoming work so the system continues serving the remainder at acceptable latency, rather than degrading for everyone.
Lock Lease Expiry concept Distributed Locking The timeout on a distributed lock that prevents a crashed holder deadlocking the system — and the source of the pattern's hardest failure mode.
Message Ordering concept Messaging & Queues The guarantee about the sequence in which messages are delivered — normally per-partition or per-group only, and lost the moment consumption is parallelised.
Message Queue Work Queue, Job Queue pattern Messaging & Queues A durable buffer between a producer and a worker that decouples them in time, absorbs bursts, and turns a synchronous dependency into a retryable one.
Message Queue tool Distributed Systems A store that holds messages until a consumer processes them, decoupling producer availability and rate from consumer availability and rate.
Metastable Failure Congestive Collapse, Death Spiral concept Failure Modes A failure state that sustains itself after the original trigger is gone, because the system's own recovery behaviour generates the load keeping it down.
Netflix: From Hystrix to Adaptive Concurrency Limits Hystrix, Concurrency Limits case-study Circuit Breakers Netflix's widely-copied circuit breaker library was retired in favour of limits that derive themselves from observed latency, because static thresholds go stale.
Offset Management concept Event Streaming How a consumer records its position in a stream, and the decision that determines whether processing is at-least-once or at-most-once.
Optimistic Claim Conditional Assignment, Compare-and-Set Dispatch pattern Consistency Models Ranking candidates from deliberately stale data and resolving the race atomically at the moment of commitment, rather than locking during selection.
Optimistic Concurrency Control OCC, Compare-and-Set pattern Distributed Locking Allowing concurrent work without locks and detecting conflict at write time by checking that the underlying version has not changed.
Outbox Pattern Transactional Outbox pattern Sagas & Compensation Writing a message to a table in the same database transaction as the state change, then publishing it asynchronously, so the two cannot diverge.
PACELC concept CAP & PACELC An extension of CAP stating that during a partition you trade availability against consistency, and else — in normal operation — you trade latency against consistency.
Partition Tolerance concept CAP & PACELC The ability to keep operating when the network drops or delays messages between nodes — not a choice, but a property of any system spanning more than one machine.
Per-Dependency Concurrency Limit Dependency Semaphore, Downstream Bulkhead pattern Bulkheads & Isolation A hard cap on in-flight calls to a specific downstream, which bounds the damage a slow dependency can do regardless of how slow it becomes - more reliable than a circuit breaker because it needs no detection.
Pivot Transaction concept Sagas & Compensation The step in a saga after which the transaction can no longer be cancelled — everything before it is compensatable, everything after it is retriable until it succeeds.
Priority Load Shedding Tiered Shedding, Criticality-Based Shedding, Selective Degradation pattern Load Shedding Dropping traffic by declared business criticality under stress rather than uniformly or randomly, exploiting the fact that the highest-volume requests are usually the least valuable ones.
Priority Queueing pattern Load Shedding Classifying requests by business importance so that overload sheds the least valuable work first rather than an arbitrary slice.
Progress-Based Liveness Are-You-Making-Progress Check, Beyond Heartbeats practice Leader Election Detecting a leader or worker that is alive but not advancing, by checking whether work is progressing rather than whether the process responds.
Quorum concept Consensus Protocols The minimum number of replicas that must respond for an operation to be considered committed, sized so that any read set and any write set must overlap.
Raft protocol Consensus Protocols A consensus algorithm designed to be understandable, decomposing agreement into leader election, log replication and safety.
Reactive Streams protocol Backpressure & Flow Control A specification for asynchronous stream processing in which the consumer requests a specific number of items, making backpressure part of the protocol rather than an afterthought.
Read Timeout Socket Timeout concept Timeouts & Deadlines The bound on how long a client waits for response data after a connection is established — distinct from the connect timeout, and the one that usually matters.
Read-Your-Writes Consistency Read-After-Write concept Consistency Models A guarantee that a client always sees its own prior writes, even in a system that is otherwise eventually consistent.
Reconnect Storm Connection Stampede concept Distributed Systems The synchronised reconnection of a large population of clients after a disruption, which routinely causes a larger outage than the disruption itself.
Redlock protocol Distributed Locking An algorithm for distributed locking across independent Redis instances, and the subject of a well-known critique about what locks can guarantee at all.
Retry Budget Retry Ratio Limit, Adaptive Retry Throttling pattern Retries & Backoff A global cap on retries as a proportion of total requests, which bounds retry amplification during broad failures in a way that per-request retry counts cannot.
Roblox 2021: A Coordination Layer as a Single Point of Failure Roblox 73-Hour Outage case-study Service Discovery A performance problem in the shared service-discovery cluster took the entire platform down for 73 hours, and its novelty made it extremely difficult to diagnose.
Saga Compensating Transaction, Long-Running Transaction pattern Distributed Transactions A sequence of local transactions across services where failure is handled by compensating actions rather than by rollback.
Saga pattern Distributed Systems A sequence of local transactions across services where each step has a compensating action that semantically undoes it if a later step fails.
Saga Orchestrator pattern Sagas & Compensation A component that explicitly drives a saga's steps and compensations, holding the flow in one place rather than distributing it across event subscriptions.
Scalability concept Distributed Systems The ability to handle growing load by adding resources, ideally with cost rising no faster than the load.
Service Discovery pattern Service Discovery How a caller finds a healthy instance of a callee in an environment where instances appear and disappear continuously.
Service Discovery concept Distributed Systems The mechanism by which a caller finds a currently healthy network address for a service whose instances are ephemeral.
Service Registry tool Service Discovery The database of currently available service instances and their addresses, maintained by registration and pruned by health checking.
Shuffle Sharding Virtual Sharding, Randomised Subset Assignment pattern Bulkheads & Isolation Assigning each tenant a random subset of the available capacity rather than a single shard, so that any two tenants rarely share their entire subset and one tenant's failure affects almost nobody completely.
Split Brain concept Leader Election A partition in which two subsets of a cluster each believe they are authoritative, accepting conflicting writes that cannot afterwards be reconciled.
Split Vote concept Leader Election An election in which no candidate obtains a majority, so the term ends with no leader and the process must repeat.
Thread Pool Isolation pattern Bulkheads & Isolation Giving each downstream dependency its own pool of threads or permits, so one slow dependency cannot consume the capacity needed to serve everything else.
Thundering Herd concept Distributed Systems A large number of clients acting simultaneously because they were synchronised by a shared event, producing a spike that the steady-state design never sized for.
Timeout Budget Deadline Propagation pattern Distributed Systems Assigning a request an overall deadline at the edge and passing the remaining time down each hop, so no service works on something already out of time.
Try-Confirm-Cancel TCC, Reservation Pattern pattern Distributed Transactions A three-phase distributed transaction where each participant first reserves resources, and a coordinator then confirms or cancels all reservations.
Two-Phase Commit 2PC, XA protocol Distributed Systems A blocking protocol for atomic commit across several resources: a coordinator asks all participants to prepare, then tells them all to commit or abort.
Version Vector Vector Clock concept Clocks & Ordering A per-replica counter set that lets a system tell whether one version causally descends from another or whether the two are genuinely concurrent.
XA Transaction protocol Distributed Transactions The X/Open standard interface for two-phase commit across heterogeneous resource managers, still common in enterprise middleware and rarely the right choice for new services.