SLO and Error Budget Service · View 05 of 21 · People and journeys
The trough, and what answers it
- The low point is phase three, where the SRE cannot distinguish an outage from an SLO that has gone blind. A platform with one alert class puts every reader in that trough on every measurement failure.
- The answer is two alert classes and a coverage floor: below 95% coverage the reliability alert is suppressed and a measurement-failure alert is raised to a different rotation instead (ADR-10).
- Phase five closes the loop by counting the night outside the page-precision figure, so a measurement problem does not quietly degrade the alerting statistics that justify the configuration.
What the architecture owes this journey
- Burn rate rather than error percentage, so the page means "the budget is genuinely going" rather than "a minute looked bad" (ADR-09).
- A short confirmation window agreeing with the long window before firing, which is what removes single-minute pages.
- no-data as a bucket state, so a silent scrape can never present as a good minute (ADR-04).
Targets
- Page precision ≥ 90% of fast-burn pages corresponding to a degradation independently recorded in the incident platform — assumed.
- Coverage floor 95% of expected buckets present; below it the verdict is insufficient-data — assumed.
- At most one page per rotation per grouped incident, remainder aggregated.