SLO and Error Budget Service · View 14 of 21 · Runtime
Decisions
- Two rules per SLO, generated from a template rather than authored per service. 1,200 hand-written rule pairs drift; generated ones can be changed estate-wide in one merge (ADR-09).
- The coverage check runs before the burn-rate comparison, not after. A blind SLO produces a measurement-failure alert and no reliability alert at all (ADR-10).
- Every evaluation is recorded with the inputs that produced it, which makes a missed page reconstructible rather than arguable — and makes page precision a measurable property of the configuration.
Targets
- Fast pair: 14.4× burn over a 1-hour long window with a 5-minute confirmation window, paging within 5 minutes (p95) — assumed.
- Slow pair: 3× burn over a 6-hour window, raising a ticket within 60 minutes (p95), never paging — assumed.
- A declared cap on simultaneous pages per grouped incident, remainder aggregated into one notification.
Risks
- Two rule pairs across 1,200 SLOs is 2,400 evaluations with their own false-positive profile. More windows would improve recall and degrade the precision that decides whether pages are still answered in six months.
- Grouping by common dependency can swallow a genuinely separate incident that happens to share one. The cap is a judgement about which failure is worse.