Term Kind Topic What it is
Aspirational and Achievable SLO practice SLI, SLO & SLA The distinction between the reliability a team wishes for and the reliability its current architecture and dependencies can actually deliver.
Availability Arithmetic concept Availability Mathematics Multiplying dependency availabilities in series and combining redundant components in parallel to derive a system's achievable availability.
Availability Calculation concept Reliability & Resilience Deriving a system's availability from its components, remembering that dependencies in series multiply.
Availability Composition concept Availability Mathematics How the availability of a system follows from its dependencies — multiplying for serial dependencies and improving sharply for redundant ones.
AWS: Static Stability Across Availability Zones Static Stability case-study Static Stability AWS designs services to keep working with the capacity they already have when a zone fails, rather than needing the control plane to provision replacements.
Blameless Postmortem Learning Review, Incident Retrospective, Non-Attributive Analysis practice Postmortems An incident analysis structured so that participants can disclose what actually happened without personal risk - which is a technical requirement for accurate information, not a cultural courtesy.
Blameless Postmortem practice Reliability & Resilience An incident review that seeks the systemic conditions that made a failure possible, explicitly excluding individual fault.
Blast Radius Failure Scope, Impact Radius, Containment Boundary concept Fault Isolation The set of users, tenants, regions or services that a single failure can affect - the quantity that distinguishes an incident from a catastrophe, and the one that architecture can most directly control.
Blast Radius Analysis practice Fault Isolation Determining, for each component, exactly what fails and who is affected when it fails completely.
Blast Radius Staging Graduated Rollout, Progressive Shedding practice Failover Applying a change or a traffic shift in increasing increments gated on health, so that a wrong decision affects a bounded population before it affects everyone.
Capacity Headroom concept Capacity Planning The deliberate gap between current load and maximum capacity, sized to absorb growth, spikes and the loss of a failure domain.
Capacity Planning practice Reliability & Resilience Deciding in advance how much capacity will be needed, given growth, seasonality and failure scenarios, and ensuring it can be there in time.
Chaos Engineering practice Reliability & Resilience Deliberately injecting failure into a system to discover, before an incident does, which of your resilience assumptions are false.
Chaos Engineering in Practice practice Chaos Engineering Deliberately injecting failure into production to verify that resilience mechanisms work — an experiment with a hypothesis, not vandalism.
Concurrency Limiting pattern Fault Isolation Bounding the number of simultaneous in-flight operations so that overload produces fast rejection rather than resource exhaustion.
Contributing Factors Analysis practice Postmortems Identifying the multiple conditions that combined to produce an incident, in place of searching for a single root cause.
Correlated Degradation Demand-Coupled Failure, Peak-Correlated Dependency Risk concept Degradation Modes The condition where a dependency's reliability worsens for the same reason that user demand rises, so the system is least capable exactly when it is most needed.
Correlated Failure Common-Mode Failure, Shared Fate concept Redundancy Failures that are not independent, so redundancy multiplies far less than the arithmetic promises - usually because replicas share code, configuration or a deployment.
Dark Failover Shadow Failover, Non-Serving Failover Rehearsal practice DR Testing Bringing a standby environment to full serving readiness and exercising it with synthetic or internal traffic while production continues untouched, so recovery is measured without an outage being risked.
Degradation Ladder Kill Switch List, Feature Shedding Order, Brownout Plan practice Degradation Modes An ordered, pre-agreed list of features that can be disabled under stress, each with a switch that works without deployment, a named owner, a trigger condition and a rehearsed fallback.
Degradation Mode concept Degradation Modes A defined, intentional reduced state of service that the system enters under specified conditions, with known behaviour and known exit criteria.
Degradation Modes concept Degradation Modes Explicitly designed operating modes below "fully working" — declared, testable, and switchable rather than emergent.
Disaster Recovery DR practice Reliability & Resilience The plan and capability for restoring service after an event that takes out a whole site, region or system.
Durable Execution Workflow Persistence, Resumable Orchestration pattern Reliability & Resilience Persisting a workflow's progress outside the process executing it, so that worker restarts, deployments and crashes resume from the last completed step rather than losing position.
Error Budget Reliability Budget, SLO Budget, Failure Allowance practice Error Budgets The permitted amount of unreliability implied by an SLO, used as an automatic arbiter between shipping features and improving reliability - and, when unspent, as evidence that the system is more reliable than …
Error Budget metric Reliability & Resilience The amount of unreliability an SLO permits, treated as a resource that feature velocity spends.
Error Budget Policy practice Error Budgets The pre-agreed consequences of exhausting an error budget, which is what turns an SLO from a metric into a decision-making mechanism.
Failover pattern Failover Switching to a standby when the primary fails — where detection, fencing and the decision to automate are harder than the switch itself.
Failover Orchestration pattern Failover The sequence of detection, decision, promotion and traffic redirection that moves service from a failed component to a healthy one.
Failover Propagation Delay Client Convergence Time, Failover Tail metric Failover The interval between a failover completing on the server side and the last client actually using the new location, which is usually the larger part of the recovery time users experience.
Failure Injection Testing practice Resilience Testing Deliberately introducing faults into a system under test to verify that timeouts, retries, fallbacks and circuit breakers behave as designed.
Fallback Independence practice Graceful Degradation The requirement that a degraded path fail independently of the primary - and be exercised continuously, because fallbacks that are never run do not work.
Fault Domain concept Fault Isolation A boundary within which a single failure is contained, defined by the infrastructure and dependencies that components inside it share.
Feature Criticality Tiering practice Graceful Degradation Classifying product functionality by whether it must work, should work, or can be dropped, so degradation decisions are made in advance by the business.
Game Day practice Reliability & Resilience A scheduled exercise in which a failure is deliberately introduced and the team responds as though it were real, to test the system and the response together.
Game Days practice Game Days Scheduled exercises where a team responds to a simulated or injected failure, testing the humans and the process as much as the system.
Google SRE: Error Budgets as a Negotiation Device SRE Error Budget case-study Error Budgets Google resolved the standing conflict between shipping speed and reliability by giving both sides a shared number and a pre-agreed consequence.
Graceful Degradation in Practice pattern Graceful Degradation Deliberately reducing functionality to preserve the core when a dependency fails — a product decision expressed in architecture.
Health Check Depth Shallow vs Deep Health Check, Dependency Health Check concept Failover How much of a service's dependency graph a health check consults, which determines whether the check reports an independent fault or a shared one that every instance will report at the same moment.
Incident Command Incident Command System, ICS practice Reliability & Resilience Assigning explicit roles during an incident — commander, operations lead, communications lead, scribe — so coordination does not compete with diagnosis.
Incident Severity Levels practice Incident Management A small, agreed scale of incident severity that determines response, escalation and communication without requiring debate during the event.
Journey-Level SLO Per-Journey Availability, User-Facing SLI practice SLI, SLO & SLA Setting reliability targets per user journey rather than per platform or per service, because consequence differs by journey and only a journey-level measurement reflects what a user experienced.
Load-Coupled Fault Injection Fault Injection Under Load, Contention Testing practice Chaos Engineering Injecting dependency faults while the system is at realistic peak load, because timeouts, pool sizes and breaker thresholds only reveal whether they are correct under contention.
N+1 and 2N Redundancy pattern Redundancy Provisioning one spare beyond required capacity versus provisioning double, and the failure assumptions each encodes.
Netflix: Chaos Monkey and the Simian Army Simian Army, Chaos Kong case-study Chaos Engineering Netflix deliberately terminated production instances during business hours to force engineers to build for failure rather than hope against it.
Nines Table concept Availability Mathematics The mapping between availability percentages and permitted downtime, and the cost curve that comes with it.
On-Call practice On-Call The rotation that responds to production problems — a system whose health is measured by whether the people in it can sustain being in it.
Page Budget Pages Per Shift, Alert Budget metric On-Call An explicit ceiling on how many pages a shift may generate, treated as a limit the team manages against rather than as an outcome it observes.
Postmortems Incident Review, Blameless Postmortem practice Postmortems Structured learning from an incident, conducted so that the truth is obtainable — which requires that telling it is safe.
Pre-Provisioned Capacity pattern Static Stability Running capacity that is already in place to absorb a failure, rather than depending on a control plane to create it during the failure.
Recovery Cost Per Unit Remediation Cost, Recoverability, Time-to-Repair Per Node concept Fault Isolation The effort required to restore one affected instance, device or tenant - the multiplier that decides whether a wide blast radius is an incident or a catastrophe, and the factor most often left out of risk asse…
Recovery Point Objective RPO metric RTO & RPO The maximum acceptable data loss measured as a duration, determining the replication and backup strategy.
Recovery Time Objective RTO metric RTO & RPO The maximum acceptable duration between a failure and restored service, agreed with the business rather than chosen by engineering.
Redundancy concept Redundancy Multiple instances of a component so that one failing does not fail the system — valuable exactly to the extent the failures are uncorrelated.
Reliability Headroom metric Capacity Planning The gap between provisioned capacity and the load that would be carried after the largest planned failure, measured under peak conditions.
Request-Based and Window-Based SLI metric SLI, SLO & SLA Two ways of computing the same reliability target — counting good events, or counting good time windows — which produce materially different numbers.
Resilience Testing practice Resilience Testing Verifying that failure-handling behaves as designed — a category of testing distinct from functional and load testing, and usually absent.
Restore Granularity Recovery Unit, Restore Scope concept RTO & RPO The smallest unit a recovery procedure can return without disturbing everything stored alongside it, which sets the real recovery time for any incident that affects a subset of the data.
Restore Verification practice DR Testing Periodically performing a real restore from backup and validating the result, as the only evidence that a recovery capability exists.
RTO and RPO metric RTO & RPO How long recovery may take, and how much data may be lost — the two numbers from which every disaster recovery design follows.