Term Kind Topic What it is
Admission Control practice Load Shedding Deciding at the entrance whether to accept a request at all, based on whether the system can complete it within its deadline.
Amazon Dynamo: Always Writeable Dynamo Paper case-study Consistency Models Amazon chose an always-writeable shopping cart with application-level conflict resolution, accepting merge complexity to guarantee that "add to cart" never fails.
Ambiguous Outcome Unknown State, Indeterminate Result concept Timeouts & Deadlines The state a system is in when a call times out - not failure, but unknown - and the design obligation to have somewhere to put it.
Atomic Commit Protocol concept Distributed Transactions Any protocol ensuring that several participants reach the same decision to commit or abort — and a problem provably unsolvable with certainty in an asynchronous system with failures.
AWS S3 2017: The Blast Radius of a Typo S3 us-east-1 Outage case-study Failure Modes A mistyped command during routine debugging removed far more capacity than intended, and the affected subsystems had not been restarted in years.
Backpressure concept Distributed Systems A mechanism by which a component under load tells its callers to slow down, rather than accepting work it cannot complete.
Backpressure in Practice Flow Control, Push-Back pattern Backpressure & Flow Control Signalling upstream to slow down so queues stay bounded, and what to do at boundaries where the producer cannot be slowed.
Bounded Queue pattern Backpressure & Flow Control A queue with a maximum depth, which converts unbounded latency growth into an explicit rejection you can control.
Brownout pattern Load Shedding Deliberately reducing the quality or completeness of every response under load, rather than serving some requests fully and rejecting others.
Bulkhead pattern Bulkheads & Isolation Partitioning resources so that exhaustion in one area cannot consume the capacity another area depends on.
Bulkheads in Practice Resource Partitioning, Compartmentalisation pattern Bulkheads & Isolation Partitioning resources so one dependency's failure cannot consume what another needs, and the aggregate-sizing mistake that moves the exhaustion one layer down.
Byzantine Fault concept Failure Modes A failure in which a component behaves arbitrarily or deceptively — returning wrong results rather than stopping — as distinct from simply crashing.
CAP Theorem Brewer's Theorem concept Distributed Systems During a network partition a distributed system must choose between consistency and availability; it cannot have both.
Cascading Timeout concept Timeouts & Deadlines The effect of independently-chosen per-hop timeouts summing to a total far longer than any caller is willing to wait.
Causal Consistency concept Consistency Models A model guaranteeing that operations which causally depend on one another are seen in the same order everywhere, while concurrent operations may be seen in any order.
Cell Isolation pattern Bulkheads & Isolation Bulkheading at the level of a complete system copy, so that any failure — including ones nobody predicted — is contained to the customers assigned to one cell.
Circuit Breaker pattern Distributed Systems A proxy that stops calling a failing dependency after a failure threshold, failing fast instead, and periodically tests whether it has recovered.
Circuit Breakers in Practice Failure Detector, Trip Switch pattern Circuit Breakers What a breaker is actually for, the configuration that stops it causing the outage it prevents, and why Netflix moved away from static thresholds.
Circular Observability Dependency Blind Spot Cycle, Self-Observing Stack, Telemetry Co-Failure concept Failure Modes When the tooling used to diagnose a failure runs on the infrastructure that is failing - so the outage removes the ability to see the outage, converting a technical problem into a search problem.
Client-Side Discovery pattern Service Discovery The caller queries the registry itself and chooses an instance, rather than sending to a stable address that something else resolves.
Clocks and Ordering Logical Clocks, Happens-Before concept Clocks & Ordering Why wall-clock timestamps cannot order events across machines, and the mechanisms that can.
Cloudflare 2019: One Regex, Global Outage Cloudflare WAF Outage case-study Load Shedding A firewall rule containing a regular expression with catastrophic backtracking consumed CPU across Cloudflare's entire global network within seconds of deployment.
Compensating Transaction pattern Sagas & Compensation A business operation that semantically undoes a previously committed step — not a rollback, because the original effect was visible and may not be fully reversible.
Compensation Impossibility Irreversible Step, Uncompensatable Action concept Sagas & Compensation The recognition that some saga steps have no true compensating action, and the design rule that follows - order steps by reversibility and gate the irreversible ones.
Conflation Last-Value Caching, Superseding Updates pattern Backpressure & Flow Control Discarding superseded updates so that a slow consumer receives the latest state rather than a backlog of stale ones - bounding work by entity count instead of by message rate.
Consistent Prefix Read concept Consistency Models A guarantee that if a sequence of writes happens in a given order, a reader sees a prefix of that sequence — never an out-of-order subset.
Consumer Group concept Event Streaming A set of consumers that cooperatively read one stream, with each partition assigned to exactly one member, so the group collectively processes every message once.
Control-Interval Matching Actuation Delay Matching, Feedback Loop Damping concept Failure Modes Setting a feedback loop's control interval to the delay of the thing it actuates - because a controller that adjusts faster than the system can respond amplifies noise instead of correcting error.
Credit-Based Flow Control protocol Backpressure & Flow Control A scheme where a receiver grants the sender a budget of bytes or messages it may transmit, replenished as the receiver consumes.
Deadline Exceeded concept Timeouts & Deadlines The error returned when a request's overall budget expires — semantically distinct from a per-hop timeout, and a signal that must not be retried blindly.
Deadline Propagation Budget Propagation, Request Deadline, Remaining-Time Passing pattern Timeouts & Deadlines Carrying the originating request's absolute deadline through every downstream hop so that each service works with the remaining budget - and refuses work that cannot finish in time, rather than starting it.
Deadline Propagation Deadline Budget, Request Deadline pattern Timeouts & Deadlines Passing the remaining time budget down each call in a request chain so downstream services never work on a request whose caller has already given up.
Distributed Locks in Practice Lease, Mutual Exclusion pattern Distributed Locking Why lease-based locking is unsafe without fencing, when locks are unavailable entirely, and the partitioning alternative that removes the problem.
Event Stream tool Distributed Systems An append-only, retained log of events that many independent consumers read at their own position, and can re-read.
Event Streaming Distributed Log, Commit Log pattern Event Streaming A durable ordered log of facts that many independent consumers read at their own pace, retaining events after consumption rather than deleting them.
Eventual Consistency concept Distributed Systems A guarantee that replicas will converge to the same value if updates stop, with no bound on how long reads may be stale.
Exponential Backoff pattern Distributed Systems Increasing the wait between retries geometrically, with random jitter, so that failures do not synchronise into a stampede.
Fail-Fast vs Fail-Safe concept Failure Modes Whether a component should stop immediately on detecting a problem, or continue in a degraded but safe mode — a choice that depends entirely on which outcome is worse.
Failure Threshold concept Circuit Breakers The condition that trips a circuit breaker — best expressed as a failure rate over a rolling window with a minimum request volume, not as a consecutive-failure count.
Fallback Strategy pattern Circuit Breakers What a caller does instead when a circuit breaker is open — the part of the pattern that determines whether failing fast helps anyone.
Fan-Out concept Distributed Systems One incoming request causing many outgoing ones, which multiplies both load and tail latency.
Fan-out on Write vs Fan-out on Read Push vs Pull Timelines concept Messaging & Queues Whether an event is copied to every recipient's store at publish time, or assembled from sources at read time - and why large systems need both.
Fault Tolerance concept Distributed Systems Continuing to operate correctly despite the failure of some components, by design rather than by luck.
Fencing Token pattern Distributed Locking A monotonically increasing number issued with a lock, checked by the resource, so a holder whose lease expired cannot act on stale authority.
GitHub 2018: 43 Seconds of Partition, 24 Hours of Recovery GitHub October 2018 Incident case-study Leader Election A 43-second network partition triggered an automated cross-region database failover, and reconciling the resulting divergence took over 24 hours.
Google Spanner: Buying Consistency With Time TrueTime, External Consistency case-study Distributed Transactions Spanner achieves globally consistent transactions by bounding clock uncertainty with dedicated hardware and deliberately waiting out that uncertainty on commit.
Graceful Degradation practice Distributed Systems Continuing to deliver reduced but useful function when a dependency fails, instead of failing the whole request.
Gray Failure Partial Failure, Fail-Slow concept Failure Modes A component that is degraded rather than down, passing health checks while serving a portion of requests slowly or incorrectly.
Grey Failure Partial Failure, Fail-Slow concept Failure Modes A component that is degraded rather than down — slow, intermittently erroring, or failing for a subset of operations — which defeats health checks built for binary states.
Half-Open State concept Circuit Breakers The circuit breaker state that allows a limited number of trial requests through to test whether a failed dependency has recovered.
Harvest and Yield concept CAP & PACELC A refinement of CAP that treats availability as a continuum — yield is the fraction of requests answered, harvest is the fraction of data reflected in an answer.
Hedged Request Request Hedging, Tied Request pattern Failure Modes Sending the same request to a second replica after a short delay and using whichever responds first, to cut tail latency caused by unlucky slow servers.
Idempotency Idempotency Key, Exactly-Once Effect pattern Idempotency The property that performing an operation twice has the same effect as performing it once — the only practical defence against the duplicates a retrying client will inevitably send.
Idempotency concept Distributed Systems The property that performing an operation many times has the same effect as performing it once.
Idempotency Key Scoping Key Scope, Deduplication Scope practice Idempotency The decision of what an idempotency key is unique within, what is stored alongside it, and how long it lives - the three choices that determine whether deduplication actually works.
Idempotency Token Store pattern Idempotency The durable record of which idempotency keys have been seen and what each one returned, and the component that decides whether the guarantee is real.
Invariant-Aligned Partitioning Partition by Invariant, Single-Writer Per Entity concept Consistency Models Choosing the partition key so that every strong-consistency invariant falls entirely inside one partition - which converts distributed coordination into local serialisation.
Jitter concept Retries & Backoff Randomising retry delays so that clients that failed together do not retry together.
Leader Election pattern Distributed Systems The process by which a group of nodes agrees which one of them is currently in charge of a task that must not run twice.
Leader Election in Practice Coordinator Election, Primary Election pattern Leader Election Selecting one node to coordinate, why majority quorum is non-negotiable, and why fencing at the resource is what actually prevents corruption.