Term Kind Topic What it is
Error Budget metric Reliability & Resilience The amount of unreliability an SLO permits, treated as a resource that feature velocity spends.
Failover Propagation Delay Client Convergence Time, Failover Tail metric Failover The interval between a failover completing on the server side and the last client actually using the new location, which is usually the larger part of the recovery time users experience.
Page Budget Pages Per Shift, Alert Budget metric On-Call An explicit ceiling on how many pages a shift may generate, treated as a limit the team manages against rather than as an outcome it observes.
Recovery Point Objective RPO metric RTO & RPO The maximum acceptable data loss measured as a duration, determining the replication and backup strategy.
Recovery Time Objective RTO metric RTO & RPO The maximum acceptable duration between a failure and restored service, agreed with the business rather than chosen by engineering.
Reliability Headroom metric Capacity Planning The gap between provisioned capacity and the load that would be carried after the largest planned failure, measured under peak conditions.
Request-Based and Window-Based SLI metric SLI, SLO & SLA Two ways of computing the same reliability target — counting good events, or counting good time windows — which produce materially different numbers.
RTO and RPO metric RTO & RPO How long recovery may take, and how much data may be lost — the two numbers from which every disaster recovery design follows.
RTO and RPO Recovery Time Objective, Recovery Point Objective metric Reliability & Resilience How long recovery may take (RTO) and how much data may be lost (RPO), the two numbers that determine the cost of a resilience design.
Service Level Indicator SLI metric Reliability & Resilience The actual measurement of a service's behaviour that an objective is set against — a ratio of good events to valid events.
Service Level Objective SLO metric Reliability & Resilience An internal target for a service level indicator, set below the level at which users notice, and used to decide whether to ship or to stabilise.
SLI, SLO and SLA metric SLI, SLO & SLA The measurement, the internal target, and the external contract — with the error budget as the mechanism that makes the target consequential.
Time to Detect MTTD metric Incident Management The interval between a problem beginning and someone knowing about it, which is often the largest and most reducible component of total incident duration.