SLO and Error Budget Service  ·  View 05 of 21  ·  People and journeys

Journey — The 02:00 Burn-Rate Page

An SRE is woken by a fast-burn alert and has to decide within a minute whether it is real.

Editable source SVG draw.io All views
Platform SRE on call, asleep Goal — Know within a minute whether this is real Trigger — A fast-burn page on the sign-in SLO Done when — Either a fixed outage or a page that should not have fired, named as such 1 · Page 14.4× burn 2 · Triage is it real? 3 · The other answer ◆ moment of truth 4 · Act fix or suppress 5 · After alert quality What they do Reads the page Checks coverage first Sees coverage 41% Not an outage — blind Fixes the scrape, not the service Logs it as measurement What the platform does Fast + confirm window agree Publishes coverage with figure Raises measurement-failure Suppresses reliability alert Keeps buckets as no-data Counts it outside precision How it feels Clear Workable Lost Where it hurts Three dashboards, three answers Did users notice or not? What answers it Burn rate, not error % Coverage on every figure Two alert classes, not one No-data is never good Precision measured per SLO Journey — 02:00, the Burn-Rate Page The trough is the moment the SRE cannot tell an outage from a blind SLO. A platform with one alert class puts every reader in that trough; two classes is the whole fix. v 1.0 · owner Reliability Architecture · date 2026-10

The trough, and what answers it

  • The low point is phase three, where the SRE cannot distinguish an outage from an SLO that has gone blind. A platform with one alert class puts every reader in that trough on every measurement failure.
  • The answer is two alert classes and a coverage floor: below 95% coverage the reliability alert is suppressed and a measurement-failure alert is raised to a different rotation instead (ADR-10).
  • Phase five closes the loop by counting the night outside the page-precision figure, so a measurement problem does not quietly degrade the alerting statistics that justify the configuration.

What the architecture owes this journey

  • Burn rate rather than error percentage, so the page means "the budget is genuinely going" rather than "a minute looked bad" (ADR-09).
  • A short confirmation window agreeing with the long window before firing, which is what removes single-minute pages.
  • no-data as a bucket state, so a silent scrape can never present as a good minute (ADR-04).

Targets

  • Page precision ≥ 90% of fast-burn pages corresponding to a degradation independently recorded in the incident platform — assumed.
  • Coverage floor 95% of expected buckets present; below it the verdict is insufficient-data — assumed.
  • At most one page per rotation per grouped incident, remainder aggregated.