Term Kind Topic What it is
Aspirational and Achievable SLO practice SLI, SLO & SLA The distinction between the reliability a team wishes for and the reliability its current architecture and dependencies can actually deliver.
Blameless Postmortem Learning Review, Incident Retrospective, Non-Attributive Analysis practice Postmortems An incident analysis structured so that participants can disclose what actually happened without personal risk - which is a technical requirement for accurate information, not a cultural courtesy.
Blameless Postmortem practice Reliability & Resilience An incident review that seeks the systemic conditions that made a failure possible, explicitly excluding individual fault.
Blast Radius Analysis practice Fault Isolation Determining, for each component, exactly what fails and who is affected when it fails completely.
Blast Radius Staging Graduated Rollout, Progressive Shedding practice Failover Applying a change or a traffic shift in increasing increments gated on health, so that a wrong decision affects a bounded population before it affects everyone.
Capacity Planning practice Reliability & Resilience Deciding in advance how much capacity will be needed, given growth, seasonality and failure scenarios, and ensuring it can be there in time.
Chaos Engineering practice Reliability & Resilience Deliberately injecting failure into a system to discover, before an incident does, which of your resilience assumptions are false.
Chaos Engineering in Practice practice Chaos Engineering Deliberately injecting failure into production to verify that resilience mechanisms work — an experiment with a hypothesis, not vandalism.
Contributing Factors Analysis practice Postmortems Identifying the multiple conditions that combined to produce an incident, in place of searching for a single root cause.
Dark Failover Shadow Failover, Non-Serving Failover Rehearsal practice DR Testing Bringing a standby environment to full serving readiness and exercising it with synthetic or internal traffic while production continues untouched, so recovery is measured without an outage being risked.
Degradation Ladder Kill Switch List, Feature Shedding Order, Brownout Plan practice Degradation Modes An ordered, pre-agreed list of features that can be disabled under stress, each with a switch that works without deployment, a named owner, a trigger condition and a rehearsed fallback.
Disaster Recovery DR practice Reliability & Resilience The plan and capability for restoring service after an event that takes out a whole site, region or system.
Error Budget Reliability Budget, SLO Budget, Failure Allowance practice Error Budgets The permitted amount of unreliability implied by an SLO, used as an automatic arbiter between shipping features and improving reliability - and, when unspent, as evidence that the system is more reliable than …
Error Budget Policy practice Error Budgets The pre-agreed consequences of exhausting an error budget, which is what turns an SLO from a metric into a decision-making mechanism.
Failure Injection Testing practice Resilience Testing Deliberately introducing faults into a system under test to verify that timeouts, retries, fallbacks and circuit breakers behave as designed.
Fallback Independence practice Graceful Degradation The requirement that a degraded path fail independently of the primary - and be exercised continuously, because fallbacks that are never run do not work.
Feature Criticality Tiering practice Graceful Degradation Classifying product functionality by whether it must work, should work, or can be dropped, so degradation decisions are made in advance by the business.
Game Day practice Reliability & Resilience A scheduled exercise in which a failure is deliberately introduced and the team responds as though it were real, to test the system and the response together.
Game Days practice Game Days Scheduled exercises where a team responds to a simulated or injected failure, testing the humans and the process as much as the system.
Incident Command Incident Command System, ICS practice Reliability & Resilience Assigning explicit roles during an incident — commander, operations lead, communications lead, scribe — so coordination does not compete with diagnosis.
Incident Severity Levels practice Incident Management A small, agreed scale of incident severity that determines response, escalation and communication without requiring debate during the event.
Journey-Level SLO Per-Journey Availability, User-Facing SLI practice SLI, SLO & SLA Setting reliability targets per user journey rather than per platform or per service, because consequence differs by journey and only a journey-level measurement reflects what a user experienced.
Load-Coupled Fault Injection Fault Injection Under Load, Contention Testing practice Chaos Engineering Injecting dependency faults while the system is at realistic peak load, because timeouts, pool sizes and breaker thresholds only reveal whether they are correct under contention.
On-Call practice On-Call The rotation that responds to production problems — a system whose health is measured by whether the people in it can sustain being in it.
Postmortems Incident Review, Blameless Postmortem practice Postmortems Structured learning from an incident, conducted so that the truth is obtainable — which requires that telling it is safe.
Resilience Testing practice Resilience Testing Verifying that failure-handling behaves as designed — a category of testing distinct from functional and load testing, and usually absent.
Restore Verification practice DR Testing Periodically performing a real restore from backup and validating the result, as the only evidence that a recovery capability exists.
Runbook Quality practice Incident Management The properties that make an operational procedure usable by a tired responder under pressure, as opposed to a document that merely exists.
Severity Levels SEV Levels practice Reliability & Resilience A predefined scale of incident impact that determines who is woken, how fast, and what process applies.
Shared-Cause Analysis What Is Shared, Redundancy Audit practice Redundancy Auditing a redundant design by asking what every replica has in common, because redundancy protects against independent failure and not against anything shared by all copies.
Sustainable On-Call practice On-Call An on-call arrangement whose alert volume, rotation size and compensation allow it to continue indefinitely without degrading the people in it.
Tabletop Exercise practice Game Days A discussion-based walkthrough of a hypothetical incident that tests procedures and decision-making without touching any system.
User Journey SLO Critical User Journey, CUJ practice SLI, SLO & SLA Defining reliability targets around what a user is trying to accomplish, measured where the user is, rather than around per-request success rates at the server.