# SLO and Error Budget Service

**Solution Architecture v1.0 · Microsoft Azure · Reliability Architecture · 2026-10 · 21 views · 17 architecture decision records**

It is Monday morning. The status page is green, the dashboards are green, and a dozen weekend tickets say checkout failed on Saturday night. The release train is loaded and someone has to decide whether it ships. In most organisations that decision is made by whoever is most senior in the room. This package is the platform that replaces the argument with arithmetic: for each critical user journey — sign in, search, add to cart, checkout, send notification, start video playback — it holds a declared reliability target, measures what users actually got, and publishes the difference as an error budget with a typed verdict attached. The assumed operating context is 1,200 SLOs across 400 services and 60 journeys, growing 30% a year, with 1.73 million minute buckets written a day and a verdict API serving 400 requests a second steady state and 2,000 for two minutes during a coordinated release window — built on Azure Data Explorer for minute-grain counters, Azure SQL as the registry and tamper-evident audit, Event Hubs for the indicator stream, Container Apps for every tier, and a Managed HSM for the one key that signs a verdict.

The product is not the number. It is **a number that can be argued with** — traceable to the forty minutes that spent it, labelled with the coverage behind it, attributable to the exact definition text in force at the time, and reproducible from retained data by anyone who disagrees.

The design rests on one rule: **the platform computes and publishes; it does not measure, and it does not enforce.** Upstream, telemetry belongs to the observability platform. Downstream, the freeze belongs to the deployment system. Everything in between is a derivation, which is why it can be recomputed, and why its blindness is describable instead of invisible.

The decisions that carry the design:

- **Counters, never ratios.** Good and valid counts are stored separately at every level. A stored ratio cannot be re-aggregated across windows and destroys the information a recompute needs — which silently removes rolling windows, exclusion overlays and interval attribution rather than merely making them expensive (ADR-01).
- **The platform owns no telemetry.** Collection stays with the observability platform. One system cannot be both the measurement path and the reliability authority, because it can no longer distinguish "I received no data" from "there was no bad traffic" (ADR-02).
- **The outcome cube is the emission contract.** Sources emit a small declared cube — status class by latency bucket — rather than a single good/valid pair, so tightening a latency threshold recomputes over 13 months of history at pre-aggregation cost. The limits of recomputability are stated rather than discovered (ADR-03).
- **Missing data is a state, not a value.** An empty bucket is `no-data`, never good and never bad; coverage is published with every figure; below the floor the verdict is `insufficient-data`. A reliability authority that produces a wrong number is worse than one that produces no number (ADR-04).
- **Definitions are versioned code.** Declarative text in the owning team's repository, reconciled into a registry that is a projection of it. A correction therefore produces a corrected history rather than a silent rewrite, and review lands at the moment a commitment changes (ADR-05).
- **Both windows published, one binds.** A trailing 28-day rolling window drives the verdict; the calendar month is sealed and reported. Two framings of the same arithmetic produce opposite incentives, so both are published and one is declared canonical (ADR-06).
- **Materialise behind a watermark.** A verdict read at 2,000 requests a second cannot scan 40,000 buckets, so budget state is an incremental projection that carries its own age — and refuses to serve past a declared staleness ceiling (ADR-07).
- **Event time, a horizon, and a key.** Buckets are assigned by event time with a 10-minute horizon, and `(slo_id, minute_utc)` is the primary key. Idempotency lives in the data model, so a replayed batch is a key collision rather than a double-count (ADR-08).
- **Page on trajectory, from a template.** Two multi-window burn-rate rules per SLO, generated rather than authored, with a confirmation window. Burn rate is the only signal with a defensible relationship to the objective, and precision is what decides whether pages are still answered in six months (ADR-09).
- **A blind SLO is a different alert.** Below the coverage floor the reliability alert is suppressed and a measurement-failure alert goes to a different rotation. One alert class leaves the on-call engineer unable to tell an outage from a broken scrape, which is the worst minute in the system (ADR-10).
- **Exclusions annotate; authors cannot approve.** An excluded interval is an overlay with a reason code, two approvers, an expiry and a permanent record, and the unexcluded figure survives forever. The privilege to remove a minute is the privilege to fake a month, so it never sits with the measured party (ADR-11).
- **Publish a signed verdict; the gate enforces.** The output is a signed, typed, time-limited document and no action at all. That keeps the authority off the critical path of everything it governs, and puts the fail-open decision where it can be reasoned about (ADR-12).
- **Fail static, and never freeze the fix.** When the platform is unreachable the gate enforces its last verified cached verdict up to a ceiling, and reliability fixes and rollbacks are exempt unconditionally. A freeze that blocks the fix for the outage that caused it is a defect (ADR-13).
- **Window close is irreversible.** At close an immutable snapshot is sealed in write-once storage. A later definition change produces a new current figure and cannot alter the record, which is what makes the record worth keeping (ADR-14).
- **Refuse the unmeasurable at authoring time.** An objective finer than minute buckets can resolve, an implicit denominator, or unbounded cardinality fails the merge. A meaningless figure published for a year costs far more than a rejected pull request (ADR-15).
- **Derived state is rebuilt, not restored.** The projection, the compiled rules and the report extracts are not backed up; their recovery objective is recompute time. One write region, because two regions deriving the same budget from overlapping buckets produce two figures no key reconciles (ADR-16).
- **Something else watches the watcher.** A minimal watchdog in a third region, sharing no store, identity or pipeline, probes the verdict API and checks its computation timestamp against an independent clock. A reliability authority cannot be the sole authority on its own availability (ADR-17).

The architecture one-pager (including why the design should still hold up in ten years, and the five risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~15 min) and [docs/decision-record.md](docs/decision-record.md) (~66 min).

---

## What is here

| Path | Contents |
|---|---|
| `diagrams/index.html` | The landing page: 21 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 17 decision records, the capability-to-technology table and the package glossary as markdown |
| `specs/part-a..c.json` | Diagram specifications, the source of truth for every view |
| `specs/manifest-a..b.json` | Acts, page titles, subtitles and reasoning cards |
| `specs/adr-onepager.json`, `specs/adr-records-a..b.json` | The one-pager, the decision records, the capability-to-technology table and the glossary |
| `scripts/build.sh` | Rebuilds every deliverable from the specs (Node only, no network) |
| `ask.md` | The requirement |

## The twenty-one views

| # | View | What it answers |
|---|---|---|
| 01 | System Context | Who declares a target, who acts on a verdict, and the two much larger things the platform refuses to own |
| 02 | High-Level Architecture | Seven stages on one spine, from an aggregate the platform did not produce to a document it does not enforce |
| 03 | Actors and Their Core Journeys | Six parties, two of them machines — including the release gate, the only actor that acts on a verdict with no human reading it |
| 04 | Journey — Can We Ship Today? | The trough is not the number: it is the forty minutes behind it, and whether they were really yours |
| 05 | Journey — The 02:00 Burn-Rate Page | The trough is the moment an SRE cannot tell an outage from an SLO that has gone blind |
| 06 | Journey — Declaring an SLO | The trough is the valid-event denominator, which is where an SLO is really defined and later quietly widened |
| 07 | Layered Architecture | Seven layers, and the seam that decides what gets backed up: below it data, above it derivation |
| 08 | Platform Components | Four tiers in one region, carrying only the three edges this view alone can show |
| 09 | Integration Surface | Three surfaces separated by their authorisation story, and an ingest path that is deliberately not an API |
| 10 | Data Flow — Cube to Verdict | How a pair of counters becomes a signed claim, and where the one irreversible step is |
| 11 | Storage Zones | Four classes, where the classification *is* the recovery design |
| 12 | Data Model | Twelve entities, and the two keys that make the arithmetic reproducible |
| 13 | Critical Flow — Asking for a Verdict | Fourteen messages, and the last one — the unreachable case — is the design |
| 14 | Burn-Rate Evaluation | From buckets to a page, with the coverage check taken ahead of every other decision |
| 15 | Recompute Lanes | Five reasons a published figure changes, and the two that cost real compute |
| 16 | Deployment Architecture | One write region across three zones, and a recovery that is a recompute rather than a failover |
| 17 | Observability | Signal type by pipeline stage, and the bottom row that exists because nothing else can see it |
| 18 | SLO Lifecycle | Eight states, and why the loop only closes at "reviewed" |
| 19 | Security Trust Zones | Five zones, and a deliberate asymmetry between reading a figure and changing one |
| 20 | Identity and Privilege Flow | Excluding forty minutes — and the one message that is refused |
| 21 | Failure Classes and Their Answers | Nine classes with their structural answer and their accepted residual, plus the two that would change the design |

## Rebuilding

```bash
bash scripts/build.sh
```

Node 20+ and nothing else. No draw.io Desktop, no browser, no network. The build assembles
`specs/views.json` and `specs/manifest.json` from the authoring parts, generates the draw.io sources,
fails on any geometry or routing defect, renders the editable and plain SVGs, writes the HTML pages
and the index, injects the one-pager and the decision record, and proves every relative link in
`diagrams/` resolves.

At v1.0 the geometry gate reports 0 errors and 0 warnings across 21 files, the routing gate 0 errors
and 0 warnings across 148 edges, all 343 nodes resolve to an icon with no weak matches, and the link
check passes on 166 relative links. The clean routing gate was reached by cutting edges rather than
by tuning them: the container, layered, trust-zone and integration views each carry only the
relationships no other view can show, and the omissions are named in each view's `note`.

## A note on the numbers

Every rate, latency, ratio, threshold and retention figure in this package is a **stated
assumption**, chosen to be defensible and arguable rather than measured. They are stated precisely
so that a reviewer can disagree with one and follow it to the decision that depends on it. `ask.md`
marks them as assumptions section by section; the decision record's evidence note restates the
operating context in one paragraph.
