Term Kind Topic What it is
Availability Arithmetic concept Availability Mathematics Multiplying dependency availabilities in series and combining redundant components in parallel to derive a system's achievable availability.
Availability Calculation concept Reliability & Resilience Deriving a system's availability from its components, remembering that dependencies in series multiply.
Availability Composition concept Availability Mathematics How the availability of a system follows from its dependencies — multiplying for serial dependencies and improving sharply for redundant ones.
Blast Radius Failure Scope, Impact Radius, Containment Boundary concept Fault Isolation The set of users, tenants, regions or services that a single failure can affect - the quantity that distinguishes an incident from a catastrophe, and the one that architecture can most directly control.
Capacity Headroom concept Capacity Planning The deliberate gap between current load and maximum capacity, sized to absorb growth, spikes and the loss of a failure domain.
Correlated Degradation Demand-Coupled Failure, Peak-Correlated Dependency Risk concept Degradation Modes The condition where a dependency's reliability worsens for the same reason that user demand rises, so the system is least capable exactly when it is most needed.
Correlated Failure Common-Mode Failure, Shared Fate concept Redundancy Failures that are not independent, so redundancy multiplies far less than the arithmetic promises - usually because replicas share code, configuration or a deployment.
Degradation Mode concept Degradation Modes A defined, intentional reduced state of service that the system enters under specified conditions, with known behaviour and known exit criteria.
Degradation Modes concept Degradation Modes Explicitly designed operating modes below "fully working" — declared, testable, and switchable rather than emergent.
Fault Domain concept Fault Isolation A boundary within which a single failure is contained, defined by the infrastructure and dependencies that components inside it share.
Health Check Depth Shallow vs Deep Health Check, Dependency Health Check concept Failover How much of a service's dependency graph a health check consults, which determines whether the check reports an independent fault or a shared one that every instance will report at the same moment.
Nines Table concept Availability Mathematics The mapping between availability percentages and permitted downtime, and the cost curve that comes with it.
Recovery Cost Per Unit Remediation Cost, Recoverability, Time-to-Repair Per Node concept Fault Isolation The effort required to restore one affected instance, device or tenant - the multiplier that decides whether a wide blast radius is an incident or a catastrophe, and the factor most often left out of risk asse…
Redundancy concept Redundancy Multiple instances of a component so that one failing does not fail the system — valuable exactly to the extent the failures are uncorrelated.
Restore Granularity Recovery Unit, Restore Scope concept RTO & RPO The smallest unit a recovery procedure can return without disturbing everything stored alongside it, which sets the real recovery time for any incident that affects a subset of the data.
Service Level Agreement SLA concept Reliability & Resilience A contractual commitment about service level, with a defined remedy — usually a service credit — when it is missed.
Static Stability Stability Under Failure concept Static Stability A system that keeps working correctly in its current state when its control plane or dependencies become unavailable, rather than needing them to stay running.
Toil concept Reliability Culture Manual, repetitive, automatable operational work that scales linearly with the size of the service and produces no lasting improvement.