Term Kind Topic What it is
RTO and RPO Recovery Time Objective, Recovery Point Objective metric Reliability & Resilience How long recovery may take (RTO) and how much data may be lost (RPO), the two numbers that determine the cost of a resilience design.
Runbook Quality practice Incident Management The properties that make an operational procedure usable by a tired responder under pressure, as opposed to a document that merely exists.
Service Level Agreement SLA concept Reliability & Resilience A contractual commitment about service level, with a defined remedy — usually a service credit — when it is missed.
Service Level Indicator SLI metric Reliability & Resilience The actual measurement of a service's behaviour that an objective is set against — a ratio of good events to valid events.
Service Level Objective SLO metric Reliability & Resilience An internal target for a service level indicator, set below the level at which users notice, and used to decide whether to ship or to stabilise.
Severity Levels SEV Levels practice Reliability & Resilience A predefined scale of incident impact that determines who is woken, how fast, and what process applies.
Shared-Cause Analysis What Is Shared, Redundancy Audit practice Redundancy Auditing a redundant design by asking what every replica has in common, because redundancy protects against independent failure and not against anything shared by all copies.
SLI, SLO and SLA metric SLI, SLO & SLA The measurement, the internal target, and the external contract — with the error budget as the mechanism that makes the target consequential.
Static Stability Stability Under Failure concept Static Stability A system that keeps working correctly in its current state when its control plane or dependencies become unavailable, rather than needing them to stay running.
Sustainable On-Call practice On-Call An on-call arrangement whose alert volume, rotation size and compensation allow it to continue indefinitely without degrading the people in it.
Tabletop Exercise practice Game Days A discussion-based walkthrough of a hypothetical incident that tests procedures and decision-making without touching any system.
Time to Detect MTTD metric Incident Management The interval between a problem beginning and someone knowing about it, which is often the largest and most reducible component of total incident duration.
Toil concept Reliability Culture Manual, repetitive, automatable operational work that scales linearly with the size of the service and produces no lasting improvement.
User Journey SLO Critical User Journey, CUJ practice SLI, SLO & SLA Defining reliability targets around what a user is trying to accomplish, measured where the user is, rather than around per-request success rates at the server.
Value-Graduated Fallback Tiered Degradation by Stake, Risk-Weighted Fallback pattern Graceful Degradation A degradation ladder whose depth depends on the value at stake, so a low-value transaction proceeds on weak signals while a high-value one fails closed - replacing a single binary decision that is always wrong…