SLO and Error Budget Service  ·  View 03 of 21  ·  People and journeys

Actors and Their Core Journeys

Six parties, two of them machines, and what each of them actually gets to do.

Editable source SVG draw.io All views
The teams being measured Service team lead 400 services Goal — I want to know whether my service is reliable enough to ship, in one number I can disagree with — and if it says no, I want to see the forty minutes that spent the budget. Core journeys Declare an SLO for a journey see view 06 Read the intervals that spent it Correct a wrong definition Release manager ships twice a day Goal — It is Monday and the train is loaded. Tell me yes or no before stand-up, and make it the same answer the SRE lead would give me. Core journeys Can we ship today? see view 04 Request an override, on the record The people who carry it Platform SRE 12 on rotation Goal — Wake me when the budget is genuinely going, not when one minute looked bad. And when the metrics pipeline breaks, tell me that instead of telling me the service is fine. Core journeys The 02:00 burn-rate page see view 05 Suppress a blind SLO Replay a quarantined batch Reliability lead owns the policy Goal — I want the freeze to be arithmetic rather than seniority, and I want to see how often we grant an exception — because a policy with a routine exception has stopped existing. Core journeys Set a budget policy Approve an exclusion window Read the override register Assurance Internal auditor quarterly Goal — Show me that last quarter's published attainment cannot have been edited after the fact, and show me who removed each excluded minute and why. Core journeys Verify a closed window Trace an exclusion to its approver Machines in the cast Release gate CI/CD, 2,000 rps peak Goal — Give me a signed verdict in under 150 ms, or tell me nothing at all — but never give me a figure you cannot stand behind, because I will act on it without a human reading it. Core journeys Fetch a verdict per deployment Fail static on a stale cache Measurement plane Monitor, Prometheus Goal — I will give you buckets, sometimes late and sometimes not at all. Do not pretend my silence was a good minute. Core journeys Emit the outcome cube Go quiet during an incident Who the Budget Is For, and What They Get to Do Person or role Journey / task External / third party Security / platform Six parties, two of them machines. The release gate is in the cast because it is the only actor that acts on a verdict without a human reading it, which is what forces the verdict to be signed and typed. v 1.0 · owner Reliability Architecture · date 2026-10

What this view settles

  • The platform has four distinct human readers with incompatible needs: a team wants attribution, a release manager wants a yes or no, an SRE wants to know whether a page is real, and an auditor wants to know the record cannot have been edited.
  • Those four needs are why the verdict is typed rather than numeric, why coverage is published alongside every figure, and why window-close snapshots have no update path.
  • The measurement plane is in the cast as an actor with a stated goal — "do not pretend my silence was a good minute" — because treating it as a reliable dependency is the single most common way this kind of platform lies.

Assumptions

  • 12 SREs on rotation, one reliability lead owning policy across the estate, quarterly internal audit.
  • Release gate traffic is machine-driven and bursty around coordinated release windows.

Risks

  • The reliability lead holds both the exclusion approval and the override grant. That concentration is deliberate for accountability and is the main reason the override register is reported weekly (ADR-11).
  • No actor here is incentivised to make an SLO harder. The admission checks and the two-approver rule are the only structural counterweight.