SLO and Error Budget Service  ·  View 14 of 21  ·  Runtime

Burn-Rate Evaluation

From buckets to a page, with one decision taken ahead of all the others.

Editable source SVG draw.io All views
Read Window reduce sum counters Coverage check floor 95% Normalise Burn rate 1x exhausts at close Evaluate Fast pair 14.4x over 1 h Slow pair 3x over 6 h Confirm Confirmation window 5 min must agree Classify Alert class reliability or measurement Grouping by dependency Route Page capped per incident Ticket never pages Measurement failure separate rotation Record Alert decision inputs retained Burn-Rate Evaluation — From Buckets to a Page Application we own Decision point External / third party Risk / gap Data store Two rules per SLO, generated from a template rather than authored, and one decision ahead of both: a blind SLO produces a measurement-failure alert and no reliability alert at all. Every evaluation is recorded with its inputs, which is what makes a missed page reconstructible instead of arguable. v 1.0 · owner Reliability Architecture · date 2026-10

Decisions

  • Two rules per SLO, generated from a template rather than authored per service. 1,200 hand-written rule pairs drift; generated ones can be changed estate-wide in one merge (ADR-09).
  • The coverage check runs before the burn-rate comparison, not after. A blind SLO produces a measurement-failure alert and no reliability alert at all (ADR-10).
  • Every evaluation is recorded with the inputs that produced it, which makes a missed page reconstructible rather than arguable — and makes page precision a measurable property of the configuration.

Targets

  • Fast pair: 14.4× burn over a 1-hour long window with a 5-minute confirmation window, paging within 5 minutes (p95) — assumed.
  • Slow pair: 3× burn over a 6-hour window, raising a ticket within 60 minutes (p95), never paging — assumed.
  • A declared cap on simultaneous pages per grouped incident, remainder aggregated into one notification.

Risks

  • Two rule pairs across 1,200 SLOs is 2,400 evaluations with their own false-positive profile. More windows would improve recall and degrade the precision that decides whether pages are still answered in six months.
  • Grouping by common dependency can swallow a genuinely separate incident that happens to share one. The cap is a judgement about which failure is worse.