Term Kind Topic What it is
Ambiguous Outcome Unknown State, Indeterminate Result concept Timeouts & Deadlines The state a system is in when a call times out - not failure, but unknown - and the design obligation to have somewhere to put it.
Atomic Commit Protocol concept Distributed Transactions Any protocol ensuring that several participants reach the same decision to commit or abort — and a problem provably unsolvable with certainty in an asynchronous system with failures.
Backpressure concept Distributed Systems A mechanism by which a component under load tells its callers to slow down, rather than accepting work it cannot complete.
Byzantine Fault concept Failure Modes A failure in which a component behaves arbitrarily or deceptively — returning wrong results rather than stopping — as distinct from simply crashing.
CAP Theorem Brewer's Theorem concept Distributed Systems During a network partition a distributed system must choose between consistency and availability; it cannot have both.
Cascading Timeout concept Timeouts & Deadlines The effect of independently-chosen per-hop timeouts summing to a total far longer than any caller is willing to wait.
Causal Consistency concept Consistency Models A model guaranteeing that operations which causally depend on one another are seen in the same order everywhere, while concurrent operations may be seen in any order.
Circular Observability Dependency Blind Spot Cycle, Self-Observing Stack, Telemetry Co-Failure concept Failure Modes When the tooling used to diagnose a failure runs on the infrastructure that is failing - so the outage removes the ability to see the outage, converting a technical problem into a search problem.
Clocks and Ordering Logical Clocks, Happens-Before concept Clocks & Ordering Why wall-clock timestamps cannot order events across machines, and the mechanisms that can.
Compensation Impossibility Irreversible Step, Uncompensatable Action concept Sagas & Compensation The recognition that some saga steps have no true compensating action, and the design rule that follows - order steps by reversibility and gate the irreversible ones.
Consistent Prefix Read concept Consistency Models A guarantee that if a sequence of writes happens in a given order, a reader sees a prefix of that sequence — never an out-of-order subset.
Consumer Group concept Event Streaming A set of consumers that cooperatively read one stream, with each partition assigned to exactly one member, so the group collectively processes every message once.
Control-Interval Matching Actuation Delay Matching, Feedback Loop Damping concept Failure Modes Setting a feedback loop's control interval to the delay of the thing it actuates - because a controller that adjusts faster than the system can respond amplifies noise instead of correcting error.
Deadline Exceeded concept Timeouts & Deadlines The error returned when a request's overall budget expires — semantically distinct from a per-hop timeout, and a signal that must not be retried blindly.
Eventual Consistency concept Distributed Systems A guarantee that replicas will converge to the same value if updates stop, with no bound on how long reads may be stale.
Fail-Fast vs Fail-Safe concept Failure Modes Whether a component should stop immediately on detecting a problem, or continue in a degraded but safe mode — a choice that depends entirely on which outcome is worse.
Failure Threshold concept Circuit Breakers The condition that trips a circuit breaker — best expressed as a failure rate over a rolling window with a minimum request volume, not as a consecutive-failure count.
Fan-Out concept Distributed Systems One incoming request causing many outgoing ones, which multiplies both load and tail latency.
Fan-out on Write vs Fan-out on Read Push vs Pull Timelines concept Messaging & Queues Whether an event is copied to every recipient's store at publish time, or assembled from sources at read time - and why large systems need both.
Fault Tolerance concept Distributed Systems Continuing to operate correctly despite the failure of some components, by design rather than by luck.
Gray Failure Partial Failure, Fail-Slow concept Failure Modes A component that is degraded rather than down, passing health checks while serving a portion of requests slowly or incorrectly.
Grey Failure Partial Failure, Fail-Slow concept Failure Modes A component that is degraded rather than down — slow, intermittently erroring, or failing for a subset of operations — which defeats health checks built for binary states.
Half-Open State concept Circuit Breakers The circuit breaker state that allows a limited number of trial requests through to test whether a failed dependency has recovered.
Harvest and Yield concept CAP & PACELC A refinement of CAP that treats availability as a continuum — yield is the fraction of requests answered, harvest is the fraction of data reflected in an answer.
Idempotency concept Distributed Systems The property that performing an operation many times has the same effect as performing it once.
Invariant-Aligned Partitioning Partition by Invariant, Single-Writer Per Entity concept Consistency Models Choosing the partition key so that every strong-consistency invariant falls entirely inside one partition - which converts distributed coordination into local serialisation.
Jitter concept Retries & Backoff Randomising retry delays so that clients that failed together do not retry together.
Lease concept Distributed Locking A lock with an expiry, granting exclusive rights for a bounded time so that a crashed holder cannot block the system forever.
Linearizability Atomic Consistency, Strong Consistency concept Consistency Models The strongest single-object guarantee — every operation appears to take effect instantaneously at some point between its call and its return.
Lock Lease Expiry concept Distributed Locking The timeout on a distributed lock that prevents a crashed holder deadlocking the system — and the source of the pattern's hardest failure mode.
Message Ordering concept Messaging & Queues The guarantee about the sequence in which messages are delivered — normally per-partition or per-group only, and lost the moment consumption is parallelised.
Metastable Failure Congestive Collapse, Death Spiral concept Failure Modes A failure state that sustains itself after the original trigger is gone, because the system's own recovery behaviour generates the load keeping it down.
Offset Management concept Event Streaming How a consumer records its position in a stream, and the decision that determines whether processing is at-least-once or at-most-once.
PACELC concept CAP & PACELC An extension of CAP stating that during a partition you trade availability against consistency, and else — in normal operation — you trade latency against consistency.
Partition Tolerance concept CAP & PACELC The ability to keep operating when the network drops or delays messages between nodes — not a choice, but a property of any system spanning more than one machine.
Pivot Transaction concept Sagas & Compensation The step in a saga after which the transaction can no longer be cancelled — everything before it is compensatable, everything after it is retriable until it succeeds.
Quorum concept Consensus Protocols The minimum number of replicas that must respond for an operation to be considered committed, sized so that any read set and any write set must overlap.
Read Timeout Socket Timeout concept Timeouts & Deadlines The bound on how long a client waits for response data after a connection is established — distinct from the connect timeout, and the one that usually matters.
Read-Your-Writes Consistency Read-After-Write concept Consistency Models A guarantee that a client always sees its own prior writes, even in a system that is otherwise eventually consistent.
Reconnect Storm Connection Stampede concept Distributed Systems The synchronised reconnection of a large population of clients after a disruption, which routinely causes a larger outage than the disruption itself.
Scalability concept Distributed Systems The ability to handle growing load by adding resources, ideally with cost rising no faster than the load.
Service Discovery concept Distributed Systems The mechanism by which a caller finds a currently healthy network address for a service whose instances are ephemeral.
Split Brain concept Leader Election A partition in which two subsets of a cluster each believe they are authoritative, accepting conflicting writes that cannot afterwards be reconciled.
Split Vote concept Leader Election An election in which no candidate obtains a majority, so the term ends with no leader and the process must repeat.
Thundering Herd concept Distributed Systems A large number of clients acting simultaneously because they were synchronised by a shared event, producing a spike that the steady-state design never sized for.
Version Vector Vector Clock concept Clocks & Ordering A per-replica counter set that lets a system tell whether one version causally descends from another or whether the two are genuinely concurrent.