Document 15 min read

Architecture One-Pager

Solution Architecture v1.0 · Microsoft Azure · Reliability Architecture · 2026-10 · 21 views · 17 architecture decision records

SLO and Error Budget Service · Solution Architecture v1.0 · Microsoft Azure · Reliability Architecture · 2026-10 · 21 views · 17 architecture decision records

The platform computes and publishes; it does not measure, and it does not enforce.

Reliability in most organisations is an argument, not a measurement. Dashboards show whatever was convenient to graph, alerts fire on error percentages that bear no defensible relationship to any commitment, and the decision to hold a release is made by whoever is most senior in the room. The underlying problem is not a missing dashboard: it is that nobody has written down what "reliable enough" means for a journey a user would recognise, in terms that can be subtracted from. An error budget does that — it converts a target into a quantity of permitted failure, which makes reliability spendable, traceable and arguable. The difficulty is that the resulting number is only worth having if it is trustworthy, and a reliability authority has far more ways to lie than to be wrong: a silent metrics pipeline reads as a perfect month, a stored ratio cannot be re-aggregated, an exclusion window in the hands of the measured party makes any breach disappear, and a figure nobody can trace to an interval gets argued with exactly like the dashboard it replaced.

Sit strictly between a measurement plane the platform does not own and an enforcement plane it does not control. Upstream, the sources emit a small, declared outcome cube — status class by latency bucket, per minute, per SLO — and the platform consumes it; it never collects, samples or stores raw signal. Internally everything is a derivation over those counters plus a versioned definition, which makes every published figure reproducible and makes a measurement outage a coverage problem rather than a reliability claim. Downstream, the platform publishes a signed, typed, time-limited verdict document and the deployment system decides what to do with it. The whole design is three stores and one rule: the registry and the minute buckets are data, everything else is a projection, and the platform would rather publish insufficient-data than a figure it cannot substantiate.

What it is, and what it is not

  • A derivation engine over someone else's aggregates — not An observability platform that happens to compute SLOs
  • A publisher of signed verdicts — not A gate that blocks deployments
  • Counters summed over a declared window — not A stored availability percentage
  • Coverage published with every figure — not A number whose confidence is left to the reader
  • no-data as a bucket state — not A quiet minute counted as a good one
  • Definitions reviewed as code, versioned immutably — not A console where targets are edited in place
  • Exclusions as annotations under two approvers — not An operator who can delete a bad afternoon
  • Window-close snapshots with no update path — not A monthly report regenerated from current state
  • Budget state as a rebuildable cache — not A derived store with a backup and an RPO

The decisions that are the architecture

  1. Counters, never ratios (ADR-01) — Good and valid counts are stored separately at every level of aggregation. A stored ratio cannot be re-aggregated across windows and destroys the information a recompute needs, which quietly makes every later decision — rolling windows, exclusion overlays, definition corrections — impossible rather than expensive.
  2. The platform owns no telemetry (ADR-02) — Collection stays with the observability platform. This is the boundary the whole set honours: it forces reproducibility, it makes a measurement outage describable rather than invisible, and it means the reliability authority is not also a hard dependency of the measurement path.
  3. The outcome cube is the emission contract (ADR-03) — Sources emit a low-cardinality cube — status class by latency bucket — rather than a single good/valid pair. Pre-aggregation is two orders of magnitude cheaper than request events, and a cube keeps retroactive redefinition possible over the declared dimensions, which is most of what teams actually want to change.
  4. Missing data is a state, not a value (ADR-04) — An empty bucket is no-data, never good and never bad. Coverage is published with every figure, and below the floor the verdict is insufficient-data. A reliability authority that produces a wrong number is worse than one that produces no number.
  5. Definitions are versioned code (ADR-05) — SLOs are declarative text in the owning team's repository, reconciled into a registry that is a projection of it. A correction therefore produces a corrected history rather than a silent rewrite, and every budget figure is attributable to the exact text that produced it.
  6. Both windows published, one drives the freeze (ADR-06) — A trailing 28-day rolling window drives the verdict; the calendar month is published alongside for reporting. Rolling makes an incident's cost decay rather than reset, which is the behaviour that matches how trust in a service actually recovers.
  7. Materialise incrementally, behind a watermark (ADR-07) — Budget state is a projection advanced by a watermark rather than recomputed per query. A verdict read at 2,000 requests a second cannot scan 40,000 buckets, and the price — a staleness window with a declared ceiling — is one the verdict can carry honestly.
  8. Event time, a horizon, and a key (ADR-08) — Buckets are assigned by event time with a 10-minute lateness horizon, and (slo_id, minute_utc) is the primary key. Idempotency lives in the data model, so a replayed batch is a key collision rather than a double-count, and late data beyond the horizon becomes a coverage deficit rather than a dropped event.
  9. Page on trajectory, from a template (ADR-09) — Two multi-window burn-rate rules per SLO, generated rather than authored, with a short confirmation window. Burn rate is the only alerting signal with a defensible relationship to the objective, and generated rules are the only way 1,200 SLOs stay consistent.
  10. A blind SLO is a different alert (ADR-10) — Below the coverage floor the reliability alert is suppressed and a measurement-failure alert goes to a different rotation. One alert class puts the on-call engineer in the position of being unable to tell an outage from a broken scrape, which is the worst minute in the whole system.
  11. Exclusions annotate; authors cannot approve (ADR-11) — An excluded interval is an overlay with a reason code, an expiry, two approvers and a permanent record, and the unexcluded figure stays retrievable forever. The privilege to remove a minute is equivalent to the privilege to fake a month, so it never sits with the measured party.
  12. Publish a signed verdict; the gate enforces (ADR-12) — The platform returns a signed, typed, time-limited document and takes no action. That keeps it off the critical path of everything it governs, and puts the fail-open or fail-closed choice in the only place that can reason about it.
  13. Fail static, and never freeze the fix (ADR-13) — When the platform is unreachable the gate enforces its last cached verdict up to a declared staleness ceiling, and changes classed as reliability fixes or rollbacks are exempt from a freeze by default. A freeze that blocks the fix for the outage that caused it is a defect.
  14. Window close is irreversible (ADR-14) — At window close an immutable snapshot is sealed in write-once storage. A later definition change or exclusion produces a new current figure and cannot alter the compliance record, which is what makes the record worth keeping.
  15. Refuse the unmeasurable at authoring time (ADR-15) — An objective whose budget is finer than minute buckets can resolve, a denominator left implicit, or label dimensions past the cardinality limit are rejected when the definition is written. A meaningless figure published for a year is far more expensive than a rejected pull request.
  16. Derived state is rebuilt, not restored (ADR-16) — Budget state, compiled rules and report extracts are not backed up; their recovery objective is recompute time. One write region, because two regions computing the same budget from overlapping buckets produce two figures no key reconciles.
  17. Something else watches the watcher (ADR-17) — A minimal watchdog in a third region probes the verdict API and holds the platform's own availability signal. A reliability authority cannot be the sole authority on its own availability, and a self-check sharing the estate's failure modes measures nothing.

Why this should still be right in ten years

Clouds, metric stores, container runtimes and alerting products will all be replaced inside the life of this design. What should survive is the set of claims about where authority lives, because those are statements about the problem rather than about Azure.

  • The boundary is not a technology. "Compute and publish, do not measure and do not enforce" is a statement about which systems may depend on which. Replace Data Explorer with something not yet shipped and the sentence is unchanged; adopt a vendor SLO product that owns collection and enforcement and the sentence is contradicted, which is the useful test to apply to any future proposal.
  • Counters outlive query engines. Storing good and valid counts separately is a consequence of arithmetic, not of a storage product. Any future engine can re-aggregate them; none can un-divide a stored ratio. This is the decision most likely to still be paying for itself in a decade, and the one most often skipped because the ratio is what people want to read.
  • Honesty about absence is permanent. Telemetry pipelines will keep breaking, in new ways, forever. A design in which absence is representable — no-data, coverage, insufficient-data — degrades gracefully under failures nobody has imagined yet. A design where absence is indistinguishable from success fails silently under all of them.
  • The enforcement seam will be tested repeatedly. Every few years someone will propose that the platform block deployments directly, because the integration is tedious and the policy feels weak. The answer does not change with scale: a reliability authority on the critical path of the release path means its outage is simultaneously an observability outage and a deployment outage.
  • Privilege design ages better than process. Separation of duties on exclusions and overrides is structural. Organisational norms about who may adjust a number will drift with every reorganisation; a rule that an author cannot approve their own exclusion survives the reorganisation, and the override register makes the drift visible when it happens anyway.
  • Rebuildability is the cheapest insurance. Treating derived state as disposable means the recovery story improves automatically as compute gets cheaper and faster, while a backup-based design's recovery story degrades as the data grows. The 90-minute recompute target gets better with time; a restore of the same volume does not.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with a number and immediately see what depends on it.

Quality Target How it is met View
Budget freshness p95 ≤ 60 s, p99 ≤ 120 s from event timestamp to updated budget state Streaming ingest with event-time bucketing, and a watermarked incremental projection rather than per-query recompute 10
Verdict latency p99 ≤ 150 ms in-region, ≤ 400 ms cross-region Constant-time read from the materialised projection, with signing on a short-lived signature rather than per request 13
Verdict availability ≥ 99.95% monthly, successful responses over valid requests per minute Zone-redundant read tier behind a global front end, with no dependency on the computation plane to serve 16
Control plane availability ≥ 99.9% monthly for authoring, registry and reporting Separate deployment from the read path, so authoring can be down while verdicts are served 08
Fast-burn detection Page dispatched within 5 minutes (p95) of a 14.4× burn beginning Multi-window evaluation on a short tick with a 5-minute confirmation window before firing 14
Slow-burn detection Ticket within 60 minutes (p95) of a sustained 3× burn Long-window pair evaluated on a slower tick, routed to a ticket queue and never to a page 14
Page precision ≥ 90% of fast-burn pages correspond to a degradation recorded in the incident platform Burn-rate rather than threshold alerting, confirmation windows, and coverage-floor suppression 17
Data coverage floor 95% of expected buckets present for a publishable figure Per-minute coverage flag, aggregated per window and published alongside every figure 10
Lateness tolerance 10-minute horizon; later data becomes a coverage deficit Event-time assignment with a bounded reopen window, then an immutable bucket 15
Single-SLO recompute ≤ 90 s (p95) over a 28-day window Columnar range scan over minute buckets partitioned by time and SLO identity 15
Full-estate backfill ≤ 6 hours, scheduled Partitioned by SLO on preemptible capacity, deployed separately from the read path 15
Ingest scale ≤ 25,000 samples/second, 1.73 million minute buckets per day at launch Aggregation upstream, so platform scale follows SLO count and not request volume 02
Estate growth 1,200 to 5,000 SLOs with no re-architecture Computation partitioned by SLO identity; cardinality bounded per definition at authoring time 12
Registry durability RPO ≤ 1 minute, RTO ≤ 15 minutes Transactional store with point-in-time restore and cross-region geo-replication 11
Derived-state recovery Trustworthy projection within 90 minutes of regional recovery Recompute from retained aggregates; derived state is not backed up at all 16
Retention Minute buckets 13 months; snapshots and audit 5 years; alert decisions 13 months Hot/warm tiering with a cold compressed archive, and write-once retention on the evidence zone 11
Reproducibility Identical inputs and definition version reproduce identical figures Counters rather than ratios, versioned immutable definitions, and deterministic window reduction 12

Scope

In scope

  • SLO definition, versioning, admission and ownership for availability, latency, freshness and correctness indicators
  • Minute-grain SLI computation from pre-aggregated upstream counters, with event-time bucketing and a declared lateness horizon
  • Data coverage measurement, and no-data as a first-class bucket state
  • Error budget accounting over rolling and calendar-aligned windows, with interval attribution
  • Immutable window-close snapshots as the compliance record
  • Multi-window burn-rate alerting, with measurement failure as a separate alert class
  • Budget policy objects, and a signed typed verdict published for the release gate to act on
  • Exclusion windows and freeze overrides under separation of duties, with a permanent record
  • Monthly reliability reporting, trend history and alert-quality measurement
  • Recompute and backfill of history after a definition correction

Explicitly out of scope

  • Collecting, sampling or storing telemetry — the observability platform owns the signal, and this platform consumes aggregates
  • Blocking, gating or stopping a deployment — the platform publishes a verdict and the deployment system acts on it
  • Dashboards, trace exploration and ad-hoc metric query, which belong to the observability platform
  • Incident management, paging rotas and escalation policy, which belong to the incident platform
  • Capacity planning and performance engineering, which consume the budget but are not computed here
  • Customer-facing SLAs, credits and contractual reporting, which are a commercial artefact derived separately
  • Deciding what a journey's target should be — the platform refuses an unmeasurable objective but never proposes one

What a four-week prototype should prove

The prototype's job is to falsify the two central claims — that a pre-aggregated outcome cube is expressive enough to be worth the cost saving, and that a derived, non-enforcing authority is actually usable by a release process. Everything else is engineering.

  1. Three real journeys on one service, with availability and latency SLOs defined as text in a repository and reconciled into a registry
  2. A genuine outcome cube emitted from existing recording rules, with its dimensions chosen before anyone knows which predicates will be wanted
  3. Minute buckets over 28 days of retained history, with good and valid counts stored separately
  4. Rolling and calendar budgets computed from the same buckets, with interval attribution down to the minute
  5. One fast and one slow burn-rate rule per SLO, generated from a template, with alert decisions recorded with their inputs
  6. A verdict API a real CI/CD pipeline calls on every deployment, with a cached fail-static path exercised deliberately
  • A predicate is redefined after the fact — from 2xx-under-800ms to 2xx-under-500ms — and the 28-day history recomputes correctly from the cube with no source change
  • A predicate is wanted that the cube cannot express, and the cost of the resulting upstream change is measured rather than assumed
  • The metrics pipeline is stopped for 40 minutes: coverage drops, the verdict turns to insufficient-data, a measurement-failure alert fires, and no reliability alert does
  • A batch is replayed twice and arrives once past the horizon: the figure is unchanged by the replay and the late batch appears as a coverage deficit
  • A 40-minute outage is injected, the budget is spent, the gate holds the release, and a rollback is correctly exempted
  • The platform is made unreachable mid-release: the gate fails static on a cached verdict, and the behaviour at the staleness ceiling is observed rather than reasoned about
  • An SLO author attempts to approve their own exclusion and is refused; a second approver completes it and both figures remain retrievable
  • A month closes, a definition error is then found and corrected, and the sealed snapshot is confirmed unchanged while the current figure moves

Open risks, carried rather than hidden

Risk If it lands Response
The outcome cube's dimensions are chosen wrongly, upstream, once A predicate teams later want cannot be expressed, and no amount of recompute recovers it. The cheapest decision in the design becomes the one that caps what the platform can ever measure. Choose the cube with the three most contested predicates already on the table, keep a per-SLO escape hatch to request event-level emission, and measure how often the hatch is used as the signal that the cube was wrong.
The release gate is bypassable in practice The freeze becomes advisory in fact while being automatic in design, which is worse than either: the arithmetic carries authority it cannot enforce, and the first bypass teaches everyone that it can be bypassed. Measure bypasses as a first-class figure next to the override register, and treat a rising bypass rate as a policy failure rather than a compliance one.
Blind minutes become routine rather than exceptional insufficient-data stops being an honest answer and becomes the normal one, at which point the derivation boundary stops paying for itself and the platform would have to own collection after all. Alert on coverage as a platform-level trend, not only per SLO, and set a threshold at which the boundary decision is formally revisited rather than quietly eroded.
Two approvers collude on exclusions Any breach can be made to look met, and the published figure becomes unfalsifiable from inside the organisation. No internal control detects this. Publish exclusion minutes and override counts in the same report as attainment, so the adjustment is always visible next to the number it improved; escalate to an external witness only if the figure ever carries contractual weight.
SLO definitions drift toward the easily met Every service passes, the budget never binds, and the platform produces defensible figures about targets nobody believes — the most common end state for this kind of system. Report attainment against objective alongside the objective's own history, so a target that was lowered is as visible as a budget that was spent, and review it at the estate level quarterly.

The reasoning behind every component and technology choice is in the Architecture Decision Record: 17 records across 5 areas, each with the alternatives that lost and what the choice costs.