SLO and Error Budget Service  ·  View 04 of 21  ·  People and journeys

Journey — Can We Ship Today?

Monday 08:40: a release manager needs a defensible answer before stand-up.

Editable source SVG draw.io All views
Release manager checkout service, Monday 08:40 Goal — A defensible yes or no before stand-up Trigger — Weekend tickets say checkout failed; the dashboards are green Done when — The train ships, or it does not, and nobody argues about which 1 · Ask open the journey 2 · See the number budget remaining 3 · Find the damage ◆ moment of truth 4 · Read the verdict ◆ moment of truth 5 · Ship or hold gate decides What they do Opens checkout journey Reads 11% remaining Clicks the bad interval Sees Sat 22:10–22:48 Reads "exhausted" Holds features, ships the fix What the platform does Resolves journey to SLOs Serves rolling + calendar Labels coverage 99.8% Attributes to intervals Links the incident record Signs the verdict Exempts rollback class How it feels Confident Fine Uneasy Where it hurts 11% of what, exactly? Was that minute really ours? What answers it Journey, not endpoint Budget in failed requests Interval links to the incident Typed verdict, one source Fixes are never frozen Journey — Monday Morning: Can We Ship Today? The trough is not the number — it is the forty minutes behind it. A budget figure nobody can trace to an interval gets argued with, which is the failure mode this view exists to prevent. v 1.0 · owner Reliability Architecture · date 2026-10

The trough, and what answers it

  • The low point is not the number — it is phase three, where the manager has to decide whether the forty minutes that spent the budget were really theirs.
  • The structural answer is interval attribution plus a link to the incident record: a budget figure that cannot be traced to an interval gets argued with, which is the exact failure this platform exists to end (ADR-01).
  • Phase five recovers because reliability fixes and rollbacks are exempt from a freeze by default. A freeze that blocks the fix for the outage that caused it is a defect, not a policy (ADR-13).

What the architecture owes this journey

  • Counters stored separately so the budget can be expressed in failed requests rather than a percentage of a percentage (ADR-01).
  • Rolling and calendar windows both published, so the manager can see which one is driving the verdict (ADR-06).
  • Coverage on the figure, so "11% remaining" is never read as more certain than the data behind it (ADR-04).

Assumptions

  • Deployments happen twice a day against a 28-day rolling window; the verdict is fetched per deployment, not per commit.
  • The incident platform already records intervals the SLO platform can correlate against.