SLO and Error Budget Service

Architecture Views

21 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

The reliability authority behind the Monday-morning argument about whether the release can ship. Read it in seven acts: what sits inside the boundary, who it is for and what they get to do, how it is put together, what it stores, what happens when a gate asks for a verdict, how it is operated, and why it is safe. The one rule the whole set honours is that this platform computes and publishes — it does not measure, and it does not enforce. Every rate, latency, threshold and retention figure here is a stated assumption, chosen to be arguable rather than measured; the package is a design, not a report on a running system.

Context and scope

What the platform is responsible for — turning someone else's aggregates into a budget and a verdict — and the two much larger things it is deliberately not responsible for, which are collecting the telemetry and stopping the release.

People and journeys

Who declares a target, who asks whether they can ship, who gets paged at two in the morning, and the three moments where this architecture becomes visible to the person living it.
03 The teams being measured Service team lead 400 services Goal — I want to know whether my service is reliable enough to ship, in one number I can disagree with — and if it says no, I want to see the forty minutes that spent the budget. Core journeys Declare an SLO for a journey see view 06 Read the intervals that spent it Correct a wrong definition Release manager ships twice a day Goal — It is Monday and the train is loaded. Tell me yes or no before stand-up, and make it the same answer the SRE lead would give me. Core journeys Can we ship today? see view 04 Request an override, on the record The people who carry it Platform SRE 12 on rotation Goal — Wake me when the budget is genuinely going, not when one minute looked bad. And when the metrics pipeline breaks, tell me that instead of telling me the service is fine. Core journeys The 02:00 burn-rate page see view 05 Suppress a blind SLO Replay a quarantined batch Reliability lead owns the policy Goal — I want the freeze to be arithmetic rather than seniority, and I want to see how often we grant an exception — because a policy with a routine exception has stopped existing. Core journeys Set a budget policy Approve an exclusion window Read the override register Assurance Internal auditor quarterly Goal — Show me that last quarter's published attainment cannot have been edited after the fact, and show me who removed each excluded minute and why. Core journeys Verify a closed window Trace an exclusion to its approver Machines in the cast Release gate CI/CD, 2,000 rps peak Goal — Give me a signed verdict in under 150 ms, or tell me nothing at all — but never give me a figure you cannot stand behind, because I will act on it without a human reading it. Core journeys Fetch a verdict per deployment Fail static on a stale cache Measurement plane Monitor, Prometheus Goal — I will give you buckets, sometimes late and sometimes not at all. Do not pretend my silence was a good minute. Core journeys Emit the outcome cube Go quiet during an incident Who the Budget Is For, and What They Get to Do Person or role Journey / task External / third party Security / platform Six parties, two of them machines. The release gate is in the cast because it is the only actor that acts on a verdict without a human reading it, which is what forces the verdict to be signed and typed. v 1.0 · owner Reliability Architecture · date 2026-10 Actors and Their Core Journeys Six parties, two of them machines, and what each of them actually gets to do. HTML page SVG draw.io

Structure

The tiers the platform is built from, the one seam between recorded and derived state, and the surfaces through which everything else reaches it.

Data

How a counter becomes a budget, which stores are authoritative, which are disposable, and the two keys that make the arithmetic reproducible.
12 service service_id PK name owning_team tier critical_journey journey_id PK name product_owner composite_rule nullable slo slo_id PK journey_id FK service_id FK owning_team current_version FK slo_version version_id PK slo_id FK indicator_type good_predicate valid_predicate objective_pct window_len, alignment effective_from label_dims bounded sli_bucket slo_id PK1 minute_utc PK2 good_count valid_count coverage ok|no-data source_class ingested_at budget_window window_id PK slo_id FK version_id FK kind rolling|calendar opened_at, closes_at watermark_minute budget_snapshot snapshot_id PK window_id FK total, consumed remaining_pct coverage_pct sealed_at immutable exclusion_window exclusion_id PK slo_id FK from_minute, to_minute reason_code approver_1, approver_2 expires_at budget_policy policy_id PK scope slo|service thresholds[] gate_mode auto|advisory unreachable open|closed exempt_classes[] verdict verdict_id PK slo_id FK state healthy|warning| exhausted|insufficient version_id FK computed_at coverage_pct signature, valid_until alert_rule rule_id PK slo_id FK burn_multiple long_window, short_window class page|ticket alert_decision decision_id PK rule_id FK evaluated_at inputs snapshot fired bool suppressed_reason N : M 1 : N 1 : N 1 : N 1 : 1 N : M 1 : N 1 : N N : M 1 : N Data Model — Definition, Bucket, Budget, Verdict Two keys carry the design. (slo_id, minute_utc) makes bucket writes idempotent, so a replayed batch cannot double-count. A sealed budget_snapshot has no update path, so a later exclusion produces a new annotation rather than a new history. budget_window carries version_id as an FK without a drawn relation, to keep the vertical runs legible; the override register and audit ledger appear in views 19 and 20. v 1.0 · owner Reliability Architecture · date 2026-10 Data Model Twelve entities, and the two keys that make the arithmetic reproducible. HTML page SVG draw.io

Runtime

What actually happens when a release gate asks a question, when a budget starts burning, and when history has to be recomputed.

Operations

Where it runs, how a region is lost and recovered, what is watched, and the loop that keeps an SLO from becoming a number nobody uses.

Assurance

Why the published figure can be trusted, which privileges could make it lie, and the nine failure classes with the residual each one leaves behind.
21 How it shows Detection Structural answer Accepted residual Telemetry gap Buckets stop arriving Coverage per SLO no-data + measurement alert Blind minutes unrecoverable Late data Bucket arrives at +12 min Lateness histogram Horizon, then coverage deficit Figure slightly pessimistic Poison input Implausible counter jump Tolerance on closed windows Quarantine, typed, replayable Coverage drops meanwhile Definition error Budget wrong, not missing Dry-run, shadow divergence Versioned rollback + recompute Verdicts served meanwhile Projection staleness Budget lags the buckets Watermark vs now Ceiling → insufficient-data Gate loses its input Verdict API down Gate cannot ask Independent watchdog Gate fails static on cache Policy weakens with age Alert storm 200 SLOs breach at once Pages per rotation Group by dependency, cap pages One real page may be grouped Regional loss Write region gone Health probes, watchdog Promote, then recompute 90 min of insufficient-data Privilege misuse Breach made to look met Exclusion minutes, override rate Two approvers, append-only Collusion defeats it Failure Classes and Their Structural Answers Nine classes, each with a residual nobody designs away. The two that would change the architecture are in the right-hand column: if blind minutes became common the platform would have to own collection after all, and if collusion on exclusions were plausible the figure would need an external witness. v 1.0 · owner Reliability Architecture · date 2026-10 Failure Classes and Their Answers Nine classes, each with the residual nobody designs away. HTML page SVG draw.io

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should falsify.

The platform computes and publishes; it does not measure, and it does not enforce.

Reliability in most organisations is an argument, not a measurement. Dashboards show whatever was convenient to graph, alerts fire on error percentages that bear no defensible relationship to any commitment, and the decision to hold a release is made by whoever is most senior in the room. The underlying problem is not a missing dashboard: it is that nobody has written down what "reliable enough" means for a journey a user would recognise, in terms that can be subtracted from. An error budget does that — it converts a target into a quantity of permitted failure, which makes reliability spendable, traceable and arguable. The difficulty is that the resulting number is only worth having if it is trustworthy, and a reliability authority has far more ways to lie than to be wrong: a silent metrics pipeline reads as a perfect month, a stored ratio cannot be re-aggregated, an exclusion window in the hands of the measured party makes any breach disappear, and a figure nobody can trace to an interval gets argued with exactly like the dashboard it replaced.

Sit strictly between a measurement plane the platform does not own and an enforcement plane it does not control. Upstream, the sources emit a small, declared outcome cube — status class by latency bucket, per minute, per SLO — and the platform consumes it; it never collects, samples or stores raw signal. Internally everything is a derivation over those counters plus a versioned definition, which makes every published figure reproducible and makes a measurement outage a coverage problem rather than a reliability claim. Downstream, the platform publishes a signed, typed, time-limited verdict document and the deployment system decides what to do with it. The whole design is three stores and one rule: the registry and the minute buckets are data, everything else is a projection, and the platform would rather publish insufficient-data than a figure it cannot substantiate.

What it is, and what it is not

A derivation engine over someone else's aggregatesAn observability platform that happens to compute SLOs
A publisher of signed verdictsA gate that blocks deployments
Counters summed over a declared windowA stored availability percentage
Coverage published with every figureA number whose confidence is left to the reader
no-data as a bucket stateA quiet minute counted as a good one
Definitions reviewed as code, versioned immutablyA console where targets are edited in place
Exclusions as annotations under two approversAn operator who can delete a bad afternoon
Window-close snapshots with no update pathA monthly report regenerated from current state
Budget state as a rebuildable cacheA derived store with a backup and an RPO

The decisions that are the architecture

01Counters, never ratios

Good and valid counts are stored separately at every level of aggregation. A stored ratio cannot be re-aggregated across windows and destroys the information a recompute needs, which quietly makes every later decision — rolling windows, exclusion overlays, definition corrections — impossible rather than expensive.

ADR-01

02The platform owns no telemetry

Collection stays with the observability platform. This is the boundary the whole set honours: it forces reproducibility, it makes a measurement outage describable rather than invisible, and it means the reliability authority is not also a hard dependency of the measurement path.

ADR-02

03The outcome cube is the emission contract

Sources emit a low-cardinality cube — status class by latency bucket — rather than a single good/valid pair. Pre-aggregation is two orders of magnitude cheaper than request events, and a cube keeps retroactive redefinition possible over the declared dimensions, which is most of what teams actually want to change.

ADR-03

04Missing data is a state, not a value

An empty bucket is no-data, never good and never bad. Coverage is published with every figure, and below the floor the verdict is insufficient-data. A reliability authority that produces a wrong number is worse than one that produces no number.

ADR-04

05Definitions are versioned code

SLOs are declarative text in the owning team's repository, reconciled into a registry that is a projection of it. A correction therefore produces a corrected history rather than a silent rewrite, and every budget figure is attributable to the exact text that produced it.

ADR-05

06Both windows published, one drives the freeze

A trailing 28-day rolling window drives the verdict; the calendar month is published alongside for reporting. Rolling makes an incident's cost decay rather than reset, which is the behaviour that matches how trust in a service actually recovers.

ADR-06

07Materialise incrementally, behind a watermark

Budget state is a projection advanced by a watermark rather than recomputed per query. A verdict read at 2,000 requests a second cannot scan 40,000 buckets, and the price — a staleness window with a declared ceiling — is one the verdict can carry honestly.

ADR-07

08Event time, a horizon, and a key

Buckets are assigned by event time with a 10-minute lateness horizon, and (slo_id, minute_utc) is the primary key. Idempotency lives in the data model, so a replayed batch is a key collision rather than a double-count, and late data beyond the horizon becomes a coverage deficit rather than a dropped event.

ADR-08

09Page on trajectory, from a template

Two multi-window burn-rate rules per SLO, generated rather than authored, with a short confirmation window. Burn rate is the only alerting signal with a defensible relationship to the objective, and generated rules are the only way 1,200 SLOs stay consistent.

ADR-09

10A blind SLO is a different alert

Below the coverage floor the reliability alert is suppressed and a measurement-failure alert goes to a different rotation. One alert class puts the on-call engineer in the position of being unable to tell an outage from a broken scrape, which is the worst minute in the whole system.

ADR-10

11Exclusions annotate; authors cannot approve

An excluded interval is an overlay with a reason code, an expiry, two approvers and a permanent record, and the unexcluded figure stays retrievable forever. The privilege to remove a minute is equivalent to the privilege to fake a month, so it never sits with the measured party.

ADR-11

12Publish a signed verdict; the gate enforces

The platform returns a signed, typed, time-limited document and takes no action. That keeps it off the critical path of everything it governs, and puts the fail-open or fail-closed choice in the only place that can reason about it.

ADR-12

13Fail static, and never freeze the fix

When the platform is unreachable the gate enforces its last cached verdict up to a declared staleness ceiling, and changes classed as reliability fixes or rollbacks are exempt from a freeze by default. A freeze that blocks the fix for the outage that caused it is a defect.

ADR-13

14Window close is irreversible

At window close an immutable snapshot is sealed in write-once storage. A later definition change or exclusion produces a new current figure and cannot alter the compliance record, which is what makes the record worth keeping.

ADR-14

15Refuse the unmeasurable at authoring time

An objective whose budget is finer than minute buckets can resolve, a denominator left implicit, or label dimensions past the cardinality limit are rejected when the definition is written. A meaningless figure published for a year is far more expensive than a rejected pull request.

ADR-15

16Derived state is rebuilt, not restored

Budget state, compiled rules and report extracts are not backed up; their recovery objective is recompute time. One write region, because two regions computing the same budget from overlapping buckets produce two figures no key reconciles.

ADR-16

17Something else watches the watcher

A minimal watchdog in a third region probes the verdict API and holds the platform's own availability signal. A reliability authority cannot be the sole authority on its own availability, and a self-check sharing the estate's failure modes measures nothing.

ADR-17

Why this should still be right in ten years

Clouds, metric stores, container runtimes and alerting products will all be replaced inside the life of this design. What should survive is the set of claims about where authority lives, because those are statements about the problem rather than about Azure.

The boundary is not a technology

"Compute and publish, do not measure and do not enforce" is a statement about which systems may depend on which. Replace Data Explorer with something not yet shipped and the sentence is unchanged; adopt a vendor SLO product that owns collection and enforcement and the sentence is contradicted, which is the useful test to apply to any future proposal.

Counters outlive query engines

Storing good and valid counts separately is a consequence of arithmetic, not of a storage product. Any future engine can re-aggregate them; none can un-divide a stored ratio. This is the decision most likely to still be paying for itself in a decade, and the one most often skipped because the ratio is what people want to read.

Honesty about absence is permanent

Telemetry pipelines will keep breaking, in new ways, forever. A design in which absence is representable — no-data, coverage, insufficient-data — degrades gracefully under failures nobody has imagined yet. A design where absence is indistinguishable from success fails silently under all of them.

The enforcement seam will be tested repeatedly

Every few years someone will propose that the platform block deployments directly, because the integration is tedious and the policy feels weak. The answer does not change with scale: a reliability authority on the critical path of the release path means its outage is simultaneously an observability outage and a deployment outage.

Privilege design ages better than process

Separation of duties on exclusions and overrides is structural. Organisational norms about who may adjust a number will drift with every reorganisation; a rule that an author cannot approve their own exclusion survives the reorganisation, and the override register makes the drift visible when it happens anyway.

Rebuildability is the cheapest insurance

Treating derived state as disposable means the recovery story improves automatically as compute gets cheaper and faster, while a backup-based design's recovery story degrades as the data grows. The 90-minute recompute target gets better with time; a restore of the same volume does not.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with a number and immediately see what depends on it.

QualityTargetHow it is metView
Budget freshness p95 ≤ 60 s, p99 ≤ 120 s from event timestamp to updated budget state Streaming ingest with event-time bucketing, and a watermarked incremental projection rather than per-query recompute 10
Verdict latency p99 ≤ 150 ms in-region, ≤ 400 ms cross-region Constant-time read from the materialised projection, with signing on a short-lived signature rather than per request 13
Verdict availability ≥ 99.95% monthly, successful responses over valid requests per minute Zone-redundant read tier behind a global front end, with no dependency on the computation plane to serve 16
Control plane availability ≥ 99.9% monthly for authoring, registry and reporting Separate deployment from the read path, so authoring can be down while verdicts are served 08
Fast-burn detection Page dispatched within 5 minutes (p95) of a 14.4× burn beginning Multi-window evaluation on a short tick with a 5-minute confirmation window before firing 14
Slow-burn detection Ticket within 60 minutes (p95) of a sustained 3× burn Long-window pair evaluated on a slower tick, routed to a ticket queue and never to a page 14
Page precision ≥ 90% of fast-burn pages correspond to a degradation recorded in the incident platform Burn-rate rather than threshold alerting, confirmation windows, and coverage-floor suppression 17
Data coverage floor 95% of expected buckets present for a publishable figure Per-minute coverage flag, aggregated per window and published alongside every figure 10
Lateness tolerance 10-minute horizon; later data becomes a coverage deficit Event-time assignment with a bounded reopen window, then an immutable bucket 15
Single-SLO recompute ≤ 90 s (p95) over a 28-day window Columnar range scan over minute buckets partitioned by time and SLO identity 15
Full-estate backfill ≤ 6 hours, scheduled Partitioned by SLO on preemptible capacity, deployed separately from the read path 15
Ingest scale ≤ 25,000 samples/second, 1.73 million minute buckets per day at launch Aggregation upstream, so platform scale follows SLO count and not request volume 02
Estate growth 1,200 to 5,000 SLOs with no re-architecture Computation partitioned by SLO identity; cardinality bounded per definition at authoring time 12
Registry durability RPO ≤ 1 minute, RTO ≤ 15 minutes Transactional store with point-in-time restore and cross-region geo-replication 11
Derived-state recovery Trustworthy projection within 90 minutes of regional recovery Recompute from retained aggregates; derived state is not backed up at all 16
Retention Minute buckets 13 months; snapshots and audit 5 years; alert decisions 13 months Hot/warm tiering with a cold compressed archive, and write-once retention on the evidence zone 11
Reproducibility Identical inputs and definition version reproduce identical figures Counters rather than ratios, versioned immutable definitions, and deterministic window reduction 12

Scope

In scope

  • SLO definition, versioning, admission and ownership for availability, latency, freshness and correctness indicators
  • Minute-grain SLI computation from pre-aggregated upstream counters, with event-time bucketing and a declared lateness horizon
  • Data coverage measurement, and no-data as a first-class bucket state
  • Error budget accounting over rolling and calendar-aligned windows, with interval attribution
  • Immutable window-close snapshots as the compliance record
  • Multi-window burn-rate alerting, with measurement failure as a separate alert class
  • Budget policy objects, and a signed typed verdict published for the release gate to act on
  • Exclusion windows and freeze overrides under separation of duties, with a permanent record
  • Monthly reliability reporting, trend history and alert-quality measurement
  • Recompute and backfill of history after a definition correction

Explicitly out of scope

  • Collecting, sampling or storing telemetry — the observability platform owns the signal, and this platform consumes aggregates
  • Blocking, gating or stopping a deployment — the platform publishes a verdict and the deployment system acts on it
  • Dashboards, trace exploration and ad-hoc metric query, which belong to the observability platform
  • Incident management, paging rotas and escalation policy, which belong to the incident platform
  • Capacity planning and performance engineering, which consume the budget but are not computed here
  • Customer-facing SLAs, credits and contractual reporting, which are a commercial artefact derived separately
  • Deciding what a journey's target should be — the platform refuses an unmeasurable objective but never proposes one

What a four-week prototype should prove

The prototype's job is to falsify the two central claims — that a pre-aggregated outcome cube is expressive enough to be worth the cost saving, and that a derived, non-enforcing authority is actually usable by a release process. Everything else is engineering.

  1. Three real journeys on one service, with availability and latency SLOs defined as text in a repository and reconciled into a registry
  2. A genuine outcome cube emitted from existing recording rules, with its dimensions chosen before anyone knows which predicates will be wanted
  3. Minute buckets over 28 days of retained history, with good and valid counts stored separately
  4. Rolling and calendar budgets computed from the same buckets, with interval attribution down to the minute
  5. One fast and one slow burn-rate rule per SLO, generated from a template, with alert decisions recorded with their inputs
  6. A verdict API a real CI/CD pipeline calls on every deployment, with a cached fail-static path exercised deliberately
  • A predicate is redefined after the fact — from 2xx-under-800ms to 2xx-under-500ms — and the 28-day history recomputes correctly from the cube with no source change
  • A predicate is wanted that the cube cannot express, and the cost of the resulting upstream change is measured rather than assumed
  • The metrics pipeline is stopped for 40 minutes: coverage drops, the verdict turns to insufficient-data, a measurement-failure alert fires, and no reliability alert does
  • A batch is replayed twice and arrives once past the horizon: the figure is unchanged by the replay and the late batch appears as a coverage deficit
  • A 40-minute outage is injected, the budget is spent, the gate holds the release, and a rollback is correctly exempted
  • The platform is made unreachable mid-release: the gate fails static on a cached verdict, and the behaviour at the staleness ceiling is observed rather than reasoned about
  • An SLO author attempts to approve their own exclusion and is refused; a second approver completes it and both figures remain retrievable
  • A month closes, a definition error is then found and corrected, and the sealed snapshot is confirmed unchanged while the current figure moves

Open risks, carried rather than hidden

RiskIf it landsResponse
The outcome cube's dimensions are chosen wrongly, upstream, once A predicate teams later want cannot be expressed, and no amount of recompute recovers it. The cheapest decision in the design becomes the one that caps what the platform can ever measure. Choose the cube with the three most contested predicates already on the table, keep a per-SLO escape hatch to request event-level emission, and measure how often the hatch is used as the signal that the cube was wrong.
The release gate is bypassable in practice The freeze becomes advisory in fact while being automatic in design, which is worse than either: the arithmetic carries authority it cannot enforce, and the first bypass teaches everyone that it can be bypassed. Measure bypasses as a first-class figure next to the override register, and treat a rising bypass rate as a policy failure rather than a compliance one.
Blind minutes become routine rather than exceptional insufficient-data stops being an honest answer and becomes the normal one, at which point the derivation boundary stops paying for itself and the platform would have to own collection after all. Alert on coverage as a platform-level trend, not only per SLO, and set a threshold at which the boundary decision is formally revisited rather than quietly eroded.
Two approvers collude on exclusions Any breach can be made to look met, and the published figure becomes unfalsifiable from inside the organisation. No internal control detects this. Publish exclusion minutes and override counts in the same report as attainment, so the adjustment is always visible next to the number it improved; escalate to an external witness only if the figure ever carries contractual weight.
SLO definitions drift toward the easily met Every service passes, the budget never binds, and the platform produces defensible figures about targets nobody believes — the most common end state for this kind of system. Report attainment against objective alongside the objective's own history, so a target that was lowered is as visible as a budget that was spent, and review it at the estate level quarterly.

SLO and Error Budget Service — Architecture One-Pager and Decision Record

Why every component and every technology on these 21 views is what it is, and what each choice costs.

The architecture behind the Monday-morning argument. The status page is green, the dashboards are green, and a dozen weekend tickets say checkout failed — and someone has to decide whether the release train ships. This platform replaces conviction with arithmetic: for each critical user journey it holds a declared reliability target, measures what users actually got, and publishes the difference as an error budget with a typed verdict attached. It is built on Microsoft Azure, and it is deliberately the thinnest useful thing in the estate: it collects no telemetry and it stops no deployments.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure in this package is a stated assumption, chosen to be defensible and arguable rather than measured. They are stated precisely so that a reviewer can disagree with one and follow it to the decision that depends on it. The requirement document marks them as assumptions section by section; the operating context assumed throughout is 1,200 SLOs across 400 services and 60 critical user journeys, growing 30% a year, with 1.73 million minute buckets written a day and a verdict API serving 400 requests a second steady state and 2,000 for two minutes during a coordinated release window.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

The derivation boundary 3

The decision that defines the architecture, and the two that follow from it directly: what the platform is allowed to own, and what it is allowed to store.

ADR-01Good and valid counts are stored separately at every level; no ratio is ever stored ADR-02The platform consumes aggregates and never collects, samples or stores telemetry ADR-03Sources emit a low-cardinality outcome cube, not a single good/valid pair

Missing data as a first-class answer 3

What the platform says when it cannot substantiate a figure, which is the difference between a reliability authority and a reassurance machine.

ADR-04An empty bucket is no-data; below the coverage floor the verdict is insufficient-data ADR-08Buckets are assigned by event time with a 10-minute horizon, and idempotency is in the primary key ADR-10Measurement failure is a separate alert class, routed to a different rotation

Definition, admission and windows 3

Where an SLO comes from, what the platform refuses to accept, and which window the freeze is computed over.

ADR-05SLO definitions are versioned immutable records reconciled from the owning team's repository ADR-06Rolling and calendar windows are both published; the rolling window drives the verdict ADR-15Admission refuses an unmeasurable objective, an implicit denominator and unbounded cardinality

Computation and publication 4

How the budget is kept fresh at read scale, how a page is decided, and the shape of the document the platform hands over.

ADR-07Budget state is materialised incrementally behind a watermark, not recomputed per query ADR-09Two multi-window burn-rate rules per SLO, generated from a template rather than authored ADR-12The platform publishes a signed, typed, time-limited verdict and takes no action ADR-13The gate fails static on its cached verdict, and reliability fixes are never frozen

Evidence, privilege and recovery 4

What makes the published record trustworthy after the fact, who is allowed to change it, and what happens when the region goes away.

ADR-11Exclusions annotate rather than delete, and the SLO's author cannot approve one ADR-14At window close an immutable snapshot is sealed and can never be altered ADR-16Derived state is rebuilt rather than restored, and there is one write region ADR-17An independent watchdog in a third region holds the platform's own availability signal

Technology by capability

Microsoft Azure was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's reliability-and-operations use cases the observability platform and health-check service are on AWS, chaos engineering is on Google Cloud, and incident management and backup are self-hosted open source — so this area had no Azure document, and reaching for the same cloud repeatedly teaches a service catalogue rather than an architecture. The second is that Azure has no first-class managed SLO and error-budget product the way Google Cloud's monitoring does, which makes it the more instructive substrate: the burn-rate engine, the window manager and the verdict signing all have to be designed rather than configured, and the decisions stay visible. Every requirement in ask.md is vendor-neutral; the table below is one defensible realisation of it, and the capability column is what a reviewer should argue with.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
SLO registry, policies, exclusions, overrides Azure SQL Database, zone-redundant, point-in-time restore, active geo-replica, ledger tables for audit Microsoft Azure PostgreSQL Flexible Server with pgaudit; Cosmos DB with strong consistency; CockroachDB self-hosted The only state needing transactions, foreign keys and a tamper-evident audit path in one engine; small enough that its durability is cheap ADR-05
SLI minute-bucket aggregates Azure Data Explorer, time-partitioned, hot/warm tiering, 13-month retention Microsoft Azure ClickHouse self-hosted; Mimir or Thanos; Synapse dedicated pools; TimescaleDB Columnar range scans over a 28-day window per SLO are the dominant read pattern, and update policies plus materialised views cover the derivation ADR-01
Indicator stream Azure Event Hubs, partitioned by SLO identity, with consumer checkpointing Microsoft Azure Kafka or Confluent Cloud; Service Bus topics; Kinesis Per-key ordering and replayable offsets, which is what makes the idempotent bucket upsert in ADR-08 safe after a consumer crash ADR-08
Budget projection Azure Cache for Redis, zone-redundant, no backup Microsoft Azure Cosmos DB single-region; in-process cache with a rebuild on start; Data Explorer materialised view read directly Constant-time reads at the verdict API's burst rate; deliberately disposable, with recovery expressed as recompute time (ADR-16) ADR-07
Computation, ingest, publication and control tiers Azure Container Apps with KEDA scaling, separate apps per tier Microsoft Azure AKS; App Service; Azure Functions; Kubernetes anywhere Independent deploy and scale per tier without operating a cluster — a reliability platform should not be the most operationally demanding thing in the estate ADR-07
Recompute and backfill Container Apps jobs on spot-priced capacity, scheduled Microsoft Azure Batch; AKS spot node pools; Databricks jobs A full-estate rebuild is interruptible and off the read path, so preemptible capacity is the right price for it ADR-16
Window snapshots and audit evidence Blob Storage with time-based immutability, RA-GRS, five-year retention Microsoft Azure S3 Object Lock; GCS bucket lock; WORM appliance Write-once with no delete path is the whole requirement in ADR-14; cross-region read access makes the evidence survive the write region ADR-14
Verdict signing Azure Key Vault Managed HSM, non-exportable signing key Microsoft Azure Key Vault standard tier; CloudHSM; an offline key with short-lived intermediates A verdict a machine acts on must be verifiable by a party that does not trust this platform's network path (ADR-12) ADR-12
Identity for humans and machines Microsoft Entra ID with SSO, MFA, and workload identity federation for the release gate Microsoft Azure Okta with OIDC federation; SPIFFE/SPIRE for workloads Four independently assignable privileges with no long-lived API keys, and OIDC federation so the gate holds no secret ADR-11
Edge, quota and throttling for the read path Azure Front Door plus API Management Microsoft Azure Application Gateway with a WAF; NGINX ingress; Cloudflare The verdict API absorbs a 5× burst during release windows and needs per-caller quota without that logic living in the service ADR-13
Upstream measurement plane Azure Monitor and Azure Monitor managed service for Prometheus, with recording rules emitting the outcome cube Microsoft Azure Self-hosted Prometheus with remote write; Datadog; Grafana Cloud Deliberately outside the platform (ADR-02); recording rules are where the outcome cube in ADR-03 is produced ADR-02
Alert routing Azure Monitor action groups, separate groups per alert class Microsoft Azure PagerDuty or Opsgenie directly; Alertmanager Two alert classes must reach two rotations by separate paths (ADR-10), and the measurement-failure path must not traverse the primary region ADR-10
Independent watchdog Azure Functions in a third region, separate subscription and alerting path Microsoft Azure An external synthetic monitoring service; a probe in another cloud; Grafana Synthetic Monitoring Small enough to be obviously correct, and sharing no store, identity or pipeline with what it watches (ADR-17) ADR-17
Reporting Scheduled extract from sealed snapshots to a reporting sink Microsoft Azure Live Power BI against Data Explorer; a static site generated per month Reports are rendered from immutable snapshots rather than current state, so a quoted figure does not move between readings (ADR-14) ADR-14

The decisions, and the alternatives that lost

The derivation boundaryThe decision that defines the architecture, and the two that follow from it directly: what the platform is allowed to own, and what it is allowed to store.

ADR-01

Good and valid counts are stored separately at every level; no ratio is ever stored

Accepted

If someone asks for last quarter's availability, can the platform compute it from what it kept?

Context
The natural thing to store is the answer: a availability percentage per minute, per hour, per day. It is what people want to read, it is what every dashboard displays, and it is one number instead of two. It is also unusable. A ratio cannot be re-aggregated: the mean of twelve hourly availability figures is not the day's availability unless every hour carried identical traffic, and traffic is never identical. Worse, a stored ratio has thrown away the denominator, so there is no way to express a budget in the only unit that is actually spendable — failed requests. Every later capability in this design depends on being able to re-reduce the same underlying counts over a different window, a different exclusion overlay or a different predicate, and all of them become impossible rather than merely expensive the moment the division happens before storage.
Decision
The unit of storage is a pair of integers per SLO per minute: good_count and valid_count, plus a coverage flag. Ratios are computed at read time, never persisted. Every window figure — rolling, calendar, attributed interval, excluded or unexcluded — is a sum of the same two columns over a different set of rows.
How it is realised on AWS
The sli_bucket table in Azure Data Explorer is keyed on (slo_id, minute_utc) with good_count, valid_count, coverage and source_class. Window reduction is a Kusto summarise over a time range; the budget is total minus consumed where both are counts. The budget projection in Redis holds the summed counters and the derived percentage together, with the counters authoritative — a reader that wants to recombine windows reads the counters, and a reader that wants to display a figure reads the percentage.
Options weighed
  • ChosenStore good and valid counts separately: Two integers per bucket instead of one float; everything downstream stays recomputable and the budget is expressible in failed requests
  • RejectedStore the ratio per bucket: Halves the storage and makes correct re-aggregation across windows impossible, which silently removes rolling windows, exclusion overlays and interval attribution
  • RejectedStore raw request events and aggregate at query time: Maximum flexibility at two orders of magnitude more storage and ingest; the platform would then own collection, contradicting ADR-02
  • RejectedStore counters plus a pre-computed ratio as the authority: Invites readers to trust the ratio and quietly reintroduces the re-aggregation bug the counters were kept to avoid
Consequences
What it buys
  • Any window, any exclusion overlay and any corrected predicate can be recomputed from retained data
  • The budget is expressible as a count of permitted failures, which is what makes it feel spendable rather than abstract
  • Interval attribution is a sort over the same two columns rather than a separate pipeline
What it costs
  • Two columns to keep consistent, and a reader who sums ratios by hand will still get a wrong answer the platform cannot prevent
  • The percentage everyone wants to read is computed on every query, so the read path carries arithmetic a stored ratio would have avoided
Choose differently when
If the platform only ever had to answer "what is the availability of this window, right now" — no rolling windows, no exclusions, no retroactive correction, no interval attribution — a stored ratio would be adequate and cheaper. In practice the first request after launch is always "what happened on Saturday", which needs the counts.
Why it holds up over time
This is arithmetic, not engineering. Any future store can re-aggregate counters and none can un-divide a ratio, so this decision should outlive every technology named in the package. It is also the decision most often skipped, because the ratio is the thing people want to see and the cost of storing it looks like a saving.
LessonIn any system that aggregates, store the operands and not the result. The result is cheap to recompute and impossible to decompose, and the capabilities you will want in six months are all decompositions.
Shown on views10 12
ADR-02

The platform consumes aggregates and never collects, samples or stores telemetry

Accepted

When the reliability authority and the measurement pipeline fail together, what is left?

Context
The commercially obvious design has the SLO platform own collection: agents in the fleet, its own time-series store, its own query language, SLOs computed over data it controls end to end. It is what several products do, it removes an integration, and it tightens the latency from event to figure. It also makes one system simultaneously the measurement path and the reliability authority, which has two consequences. Operationally, its outage is both an observability outage and a loss of the signal that would tell you an observability outage is happening. Epistemically, it can no longer distinguish "I did not receive data" from "there was no bad traffic", because it is the thing that would have received it — and a platform that cannot represent its own blindness will report blindness as health.
Decision
Collection belongs to the observability platform. This platform's inputs are pre-aggregated counters arriving on a stream, and its outputs are derived figures and signed verdicts. It retains aggregates and derivations; it retains no raw signal, no traces and no logs of the measured systems. Enforcement, symmetrically, belongs to the deployment system (ADR-12).
How it is realised on AWS
Azure Monitor and Azure Monitor managed service for Prometheus remain the system of record for telemetry. Recording rules in the measurement plane emit the outcome cube to Event Hubs; the platform's bucketer consumes it. Nothing in the platform scrapes, and no agent is deployed to a measured workload. The layered view draws the measurement plane as the bottom band, labelled "outside", so no reader of the set mistakes it for a component.
Options weighed
  • ChosenConsume aggregates; own no collection: One integration and a dependency on someone else's uptime, in exchange for reproducibility and the ability to describe blindness
  • RejectedOwn collection end to end: Tighter latency and no integration; makes the reliability authority a hard dependency of the measurement path and unable to represent its own blindness
  • RejectedConsume aggregates but keep a parallel collection path for critical journeys: Two sources of truth for the same SLO, and the first disagreement between them is unresolvable
  • RejectedBuy a managed SLO product that owns both collection and enforcement: Fastest to a figure; concentrates measurement, derivation and enforcement in one vendor outage
Consequences
What it buys
  • A measurement outage is describable as a coverage deficit rather than invisible as a good month
  • Every published figure is a reproducible function of retained aggregates plus a versioned definition
  • The platform stays small enough that it is not the most operationally demanding thing in the estate
What it costs
  • The platform cannot fix bad telemetry; it can only refuse to launder it, which moves some problems to a team that does not own the budget
  • The outcome cube's dimensions are chosen in a system this platform does not control (ADR-03)
  • Event-to-figure latency carries an extra hop, which is why budget freshness is specified at 60 s rather than seconds
Choose differently when
If blind minutes became routine rather than exceptional — coverage regularly below the floor across many SLOs — the boundary would stop paying for itself, because a platform that answers insufficient-data most of the time is not an authority. At that point owning collection for the critical journeys would be the correct response, and the decision should be formally revisited rather than quietly eroded.
Why it holds up over time
"Compute and publish, do not measure and do not enforce" is a statement about which systems may depend on which, not about any product. It is also the single most useful test to apply to future proposals: a vendor platform that owns collection and gating contradicts it, and that contradiction is the thing to argue about rather than the feature list.
LessonDo not let the system that judges a thing also be the system that observes it and the system that punishes it. Separating the three is what makes each one's failure describable by the others.
Shown on views01 07 21
ADR-03

Sources emit a low-cardinality outcome cube, not a single good/valid pair

Accepted

When a team wants to change what "good" means, how much history can they take with them?

Context
Having decided to consume aggregates (ADR-02), the obvious contract is the one the platform needs: each source emits good_count and valid_count per minute per SLO. It is minimal and cheap. It also freezes the good-event predicate at the point of emission, which means the first time a team wants to tighten a latency threshold from 800 ms to 500 ms, there is no history to recompute against — the distinction was discarded upstream before the platform ever saw it. Request-level events would preserve everything, at two orders of magnitude more ingest and storage, and would drag the platform back across the collection boundary. The middle position is to notice that almost every predicate change teams actually want moves along one of two axes: which status classes count as failures, and which latency threshold counts as slow.
Decision
Sources emit a small declared cube per SLO per minute — counts bucketed by status class and by latency bucket — rather than a single good/valid pair. The good-event predicate becomes a projection over the cube's declared dimensions, so a predicate change along those dimensions is recomputable over retained history at pre-aggregation cost. A change outside the declared dimensions is not recomputable, and the platform says so rather than pretending.
How it is realised on AWS
Prometheus recording rules in the measurement plane emit a histogram-shaped record: for each SLO and minute, counts keyed by status class (2xx, 4xx, 5xx, client-cancelled) and by latency bucket on a fixed ladder. The bucketer flattens the cube into sli_bucket rows keyed by (slo_id, minute_utc) with the cube retained as a nested column, so window reduction sums the cells the current predicate selects. Cardinality is bounded by the admission check in ADR-15: the cube's cell count per SLO per minute is capped, which is what keeps this affordable at 1.73 million buckets a day.
Options weighed
  • ChosenA small declared outcome cube: A few cells per bucket instead of two integers; keeps retroactive redefinition possible along the axes teams actually use
  • RejectedA single good/valid pair: Cheapest contract; makes every predicate change forward-only, which is the complaint that arrives in month two
  • RejectedRequest-level events into the platform: Any predicate, any time, full attributability; two orders of magnitude more cost and it crosses the collection boundary (ADR-02)
  • RejectedCube for critical journeys, pair for the long tail: Tempting on cost; produces two classes of SLO with different capabilities and a migration whenever a journey is promoted
Consequences
What it buys
  • Tightening a latency threshold or reclassifying a status code recomputes over 13 months of history with no upstream change
  • Cost stays within a small multiple of the minimal contract, because the cube is bounded rather than open
  • The limits of recomputability are explicit — a team can be told which changes are free and which need a source change
What it costs
  • The cube's dimensions are chosen once, upstream, by a team that does not own this platform, and a dimension nobody thought to emit is unrecoverable
  • A nested column per bucket is more complex to reduce than two integers, and the reduction has to stay deterministic for ADR-01 to hold
  • Teams will occasionally want a predicate the cube cannot express, and the answer is a genuine upstream change rather than a configuration one
Choose differently when
If storage and ingest for request-level events became cheap enough that 2.5 million events a second could be retained for 13 months without a separate budget conversation, the cube would be an unnecessary compromise and event-level ingest would dominate it — though the collection-boundary objection in ADR-02 would still stand, so the events would have to arrive from the measurement plane rather than be collected here.
Why it holds up over time
The cube is a bet about which questions get asked, not about a technology. The specific dimensions will be revised; the principle — pre-aggregate along the axes along which the definition is likely to move, and declare the limits — transfers to any pre-aggregation decision in any system.
LessonWhen you pre-aggregate, you are choosing which future questions remain answerable. Spend the extra dimension on the axis your definitions will move along, and say plainly which questions you have made unanswerable.
Shown on views06 10

Missing data as a first-class answerWhat the platform says when it cannot substantiate a figure, which is the difference between a reliability authority and a reassurance machine.

ADR-04

An empty bucket is no-data; below the coverage floor the verdict is insufficient-data

Accepted

What does the platform publish when it did not receive the data?

Context
A minute with zero valid events is ambiguous: it may be three in the morning on a low-traffic journey, or the metrics pipeline may have stopped. The convenient implementations are both wrong. Treating the minute as good — a ratio of 0/0 rendered as 100% — means a stopped scrape presents as a perfect month, which is the single most common way this class of platform lies. Treating it as bad means every quiet night burns budget and the SLO becomes unusable for anything but the busiest journeys. Interpolating from neighbouring minutes is worse than either, because it manufactures evidence. The deeper point is that a reliability authority has an asymmetric cost function: a wrong figure is materially worse than no figure, because a wrong figure is acted on.
Decision
A bucket with no valid events is recorded in a third state, no-data, which is neither good nor bad and is excluded from both numerator and denominator. Coverage — the fraction of expected buckets present — is computed per window and published alongside every figure. Below a declared coverage floor the verdict is the typed value insufficient-data rather than a number, and reliability alerts for that SLO are suppressed in favour of a measurement-failure alert (ADR-10).
How it is realised on AWS
sli_bucket carries a coverage column with values ok and no-data. Window reduction counts expected minutes against present-and-ok minutes to produce coverage_pct, which is written into the budget projection and returned in every verdict and query response. The verdict API evaluates coverage before anything else and short-circuits to insufficient-data below 95%. Low-traffic journeys declare an aggregation rule — a minimum valid-event count per evaluation period — rather than being silently imputed.
Options weighed
  • Chosenno-data as a third state, with a published coverage floor: Costs a column, a floor nobody can derive from first principles, and a fourth verdict value; makes blindness representable
  • RejectedTreat an empty bucket as good (0/0 = 100%): Simplest and the most dangerous: a stopped pipeline is indistinguishable from a perfect month
  • RejectedTreat an empty bucket as bad: Fails safe in principle; makes every low-traffic journey permanently in breach and the platform unusable for the long tail
  • RejectedInterpolate from neighbouring minutes: Produces a plausible figure with no evidence behind it, which is the one outcome worse than no figure
Consequences
What it buys
  • A metrics outage is visible as itself rather than as good news
  • Every figure carries its own confidence, so "11% remaining" is never read as more certain than the data behind it
  • Regional recovery has an honest intermediate state: insufficient-data for 90 minutes while the projection rebuilds (ADR-16)
What it costs
  • The 95% floor is a judgement with no principled derivation, and SLOs will sit just above and just below it for reasons unrelated to reliability
  • insufficient-data is a fourth state every consumer has to handle, and a consumer that treats it as healthy has reintroduced the original bug
  • A low-traffic journey needs an explicit aggregation rule, which is extra work at definition time
Choose differently when
If every SLO in the estate were on a journey with continuous high traffic and a measurement pipeline with its own hard availability guarantee, the ambiguity would largely disappear and a two-state bucket would be adequate. That is not true of the long tail of 400 services, and it is never true during an incident — which is exactly when the figure is read.
Why it holds up over time
Telemetry pipelines will keep breaking in new ways indefinitely. A design in which absence is representable degrades gracefully under failures nobody has imagined yet; one in which absence is indistinguishable from success fails silently under all of them. This should be among the last decisions anyone revisits.
LessonGive every measurement system an explicit way to say "I do not know", and make the consumers handle it. A system that can only report good or bad will report its own blindness as good.
Shown on views05 10 17 21
ADR-08

Buckets are assigned by event time with a 10-minute horizon, and idempotency is in the primary key

Accepted

What happens when the same minute's data arrives twice, and when it arrives too late?

Context
Aggregates arrive over a stream, which means they arrive at-least-once, out of order, and sometimes long after the minute they describe. Assigning them to buckets by arrival time is simpler and wrong: a backlog drained after a pipeline incident would attribute an hour of old failures to the minute the backlog cleared, which both hides the real interval and manufactures a fake one. Event-time assignment is correct but leaves two questions: how long a bucket stays open for late arrivals, and what stops a replayed batch from double-counting. An unbounded reopen window means no figure is ever final; an exactly-once delivery guarantee is not available from any stream the platform would use.
Decision
Every aggregate is assigned to its bucket by event time. A bucket accepts late arrivals for 10 minutes after its close and is then immutable; data arriving later is recorded as a coverage deficit rather than silently dropped or silently applied. Idempotency is structural: (slo_id, minute_utc) is the bucket's primary key, so a replayed batch is a key collision resolved by overwrite-with-identical-value rather than an addition.
How it is realised on AWS
The bucketer consumes Event Hubs partitions with checkpointing, derives minute_utc from the record's own timestamp, and upserts sli_bucket rows on (slo_id, minute_utc) — the write is the full cell set for that minute from that source, not a delta, which is what makes a replay a no-op. A late record whose bucket has passed the horizon increments the window's coverage deficit. An input batch that would move a closed window's figure beyond a declared tolerance is quarantined with a typed reason rather than applied, and the quarantine is a retained, replayable store rather than a dead-letter queue.
Options weighed
  • ChosenEvent time, 10-minute horizon, idempotent upsert on (slo_id, minute): Correct attribution, bounded finality, and double-counting made impossible by the key rather than by the consumer's care
  • RejectedArrival-time bucketing: Trivial to implement; a drained backlog invents an incident and hides the real one
  • RejectedEvent time with an unbounded reopen window: Never loses a record; no figure is ever final, so nothing can be sealed and the compliance record becomes provisional forever
  • RejectedAdditive deltas with a de-duplication table: Standard streaming pattern; replaces a primary key with a second system that has to be correct, and its failure mode is a silently wrong count
Consequences
What it buys
  • An interval attribution points at the minutes users actually suffered, which is what makes view 04's trough answerable
  • A replay — after a bucketer crash, a source backfill or an operator retry — cannot change a figure
  • Windows become final, so snapshots can be sealed and the compliance record means something (ADR-14)
What it costs
  • Data later than 10 minutes is lost as measurement and visible only as reduced coverage, which makes the figure slightly pessimistic
  • The upsert must carry the full cell set per source per minute, so the write is larger than a delta
  • Multiple sources writing the same SLO's minute need source_class in the key path or they overwrite each other
Choose differently when
If the upstream pipeline offered a hard bound on delivery latency — every record within 30 seconds or not at all — the horizon could shrink to near zero and the coverage-deficit path would be mostly dead code. Conversely, if the platform ever had to accept bulk historical imports as a normal operation, the horizon would have to become a per-source declaration rather than a constant.
Why it holds up over time
Event-time semantics and key-based idempotency are the two things stream processing has converged on for good reasons, and neither depends on a product. The horizon's specific length is a tuning parameter that will change; the existence of a declared horizon, after which a figure is final, is what makes everything downstream possible.
LessonPut idempotency in the primary key, not in the consumer's logic. A duplicate that collides is a non-event; a duplicate that has to be detected is a bug waiting for load.
Shown on views10 12 15
ADR-10

Measurement failure is a separate alert class, routed to a different rotation

Accepted

When the SLO goes blind, who is woken, and what are they told?

Context
Once blindness is representable (ADR-04), the alerting question follows: what happens when coverage collapses? Doing nothing is unacceptable — a silently blind SLO is an unmonitored journey. Firing the reliability alert is worse than nothing, because it tells the on-call engineer that users are suffering when the truth is that the platform cannot see. That is the single worst minute in the whole system: the engineer is awake, under time pressure, and holding a signal that points at the wrong system entirely. The diagnosis they need — "this is a metrics problem, not a service problem" — is a different conclusion, reached by a different person, with a different runbook.
Decision
Below the declared coverage floor, reliability alerts for that SLO are suppressed and a distinct measurement-failure alert is raised instead, routed to the platform rotation rather than the service rotation. The two classes are never merged, and the measurement-failure alert names the SLO, the coverage figure and the source class that stopped.
How it is realised on AWS
The burn-rate evaluator checks coverage before evaluating any rule and emits one of two alert types. Measurement-failure alerts route through a separate Azure Monitor action group to the platform rotation; reliability alerts route to the owning service's rotation. The suppression is recorded on the alert_decision row with suppressed_reason, so a suppressed window is explicable afterwards and does not silently vanish. Nights attributed to measurement failure are excluded from the page-precision figure in ADR-09, so a broken pipeline cannot degrade the statistics that justify the alerting configuration.
Options weighed
  • ChosenTwo alert classes, separate rotations, with suppression recorded: An extra alert type and routing to maintain; the engineer is told which system is broken
  • RejectedOne alert class covering both: Simplest routing; puts every reader in the position of being unable to tell an outage from a blind SLO
  • RejectedSuppress silently below the floor: No false pages at all; an unmonitored journey nobody is told about, which is the failure this platform exists to prevent
  • RejectedDegrade the reliability alert's severity instead of reclassifying: One mechanism, less plumbing; a low-severity reliability alert still points the reader at the wrong system
Consequences
What it buys
  • The page names the broken system, so triage starts in the right place
  • Pipeline health becomes a platform concern with its own rotation rather than an unexplained service symptom
  • Alert precision stays a meaningful figure, because measurement failures do not pollute it
What it costs
  • Two rotations must both exist and both be staffed, which is an organisational dependency this design assumes
  • An SLO can sit blind and un-alerted to its owning team, who find out from the coverage figure rather than from a page
  • A partial measurement failure above the floor still produces a reliability alert computed on thin data
Choose differently when
If one team owned both the measurement pipeline and the services being measured, the split would be bureaucratic rather than useful and a single rotation with a well-worded alert body would be adequate. The split exists because the two systems have different owners.
Why it holds up over time
The distinction between "the thing is broken" and "my view of the thing is broken" is permanent in any monitoring system. Products and routing will change; the requirement that those two conclusions reach different people with different runbooks will not.
LessonNever let a monitoring system report its own blindness through the same channel it reports failures. The reader cannot distinguish them, and they will act on the wrong one under time pressure.
Shown on views05 14 17 21

Definition, admission and windowsWhere an SLO comes from, what the platform refuses to accept, and which window the freeze is computed over.

ADR-05

SLO definitions are versioned immutable records reconciled from the owning team's repository

Accepted

When a target turns out to have been defined wrongly for three months, what happens to those three months?

Context
A console where SLOs are edited in place is the fastest thing to build and the easiest to use. It also makes a definition change indistinguishable from a definition correction, destroys the association between a published figure and the text that produced it, and leaves no review step at the one moment when review matters most — because an SLO is a commitment, and widening a denominator is a commitment change dressed as a configuration tweak. The alternative is to treat the definition as code: text in the owning team's repository, reviewed by pull request, reconciled into a registry that is a projection of it rather than a system of authorship.
Decision
SLO definitions are declarative text in the owning team's repository. A reconciler reads merged definitions and writes versioned, immutable records into the registry, each with an effective-from timestamp. Prior versions remain readable forever, so every historical figure stays attributable to the exact text that produced it, and a correction produces a corrected history rather than a silent rewrite.
How it is realised on AWS
A Container Apps job polls the definitions repository and reconciles merged changes into Azure SQL, writing a new slo_version row rather than updating one. slo carries current_version as an FK; every budget_window and verdict records the version_id it was computed under. The admission checks in ADR-15 run in the reconciler, so a definition that cannot be measured fails the merge rather than entering the registry. The console is read-mostly: it can pause an SLO and request an exclusion, but it cannot author a predicate.
Options weighed
  • ChosenRepository-authored text, reconciled into a versioned registry: Review at the right moment, permanent attribution, and corrections that are visible as corrections; slower to change than a console
  • RejectedConsole editing with an audit log: Fast and comfortable; the audit log records that a change happened without making the prior definition's figures recomputable
  • RejectedVersioned records authored through an API, no repository: Keeps immutability and loses the review step, which is most of the value — nobody reviews an API call
  • RejectedRepository as the only store, no registry: Purest GitOps; makes every verdict read a repository read, and the registry's transactional guarantees are needed for exclusions and policy
Consequences
What it buys
  • A wrong definition is corrected and the history recomputes, instead of the error being frozen into the record
  • Widening a denominator is a reviewed change with a named author, which is the only structural counterweight to targets drifting easier
  • Every figure in a monthly report can be traced to the text in force at the time
What it costs
  • Changing an SLO takes a pull request, which teams will find slow in exactly the moment they want it to be fast
  • The registry can lag the repository, so there is a reconciliation window during which the two disagree
  • Definition ownership stays with the service team, which makes targets achievable and comparability across 400 services weak
Choose differently when
If SLO definitions needed to change inside minutes — an incident-time adjustment, a dynamically generated SLO per tenant — a repository round trip would be the wrong mechanism and an API with versioned records would dominate. Nothing in this problem has that shape: an SLO that changes hourly is not a commitment.
Why it holds up over time
"The definition is reviewed text and the registry is its projection" survives any change of version-control system or registry technology. The specific reconciler is disposable; the claim that authorship lives outside the serving system is what keeps corrections honest.
LessonIf a change to a definition changes the meaning of history, make the definition immutable and versioned, and put the review where the commitment is made rather than where the data is served.
Shown on views06 12 18
ADR-06

Rolling and calendar windows are both published; the rolling window drives the verdict

Accepted

Should the cost of Saturday's outage decay gradually, or reset on the first of the month?

Context
The two window shapes produce opposite incentives over identical arithmetic. A trailing 28-day rolling window never resets: an incident's cost decays minute by minute, which matches how trust in a service actually recovers, and means a team can be frozen for four weeks over one bad Saturday. A calendar month resets on a date: it is comprehensible, it aligns with every other report in the business, and it rewards holding risky changes until the first while making the last week of a bad month unshippable. Both can be computed from the same counters (ADR-01) and both can be published. Only one can drive the freeze, because two verdicts for the same SLO is not a verdict.
Decision
Both windows are computed and published for every SLO. The trailing 28-day rolling window is canonical for the verdict and therefore for the freeze; the calendar month is the reporting and compliance window and is the one sealed at close (ADR-14). Every response states which window produced the figure it carries.
How it is realised on AWS
budget_window rows carry kind as rolling or calendar with their own open and close timestamps, both reducing over the same sli_bucket rows. The budget projection materialises both. The verdict API reads the rolling figure and names it explicitly in the response; the monthly report and the sealed snapshot use the calendar window. The console shows the two side by side, because a release manager asking "why does it say exhausted" needs to see which window is binding.
Options weighed
  • ChosenPublish both; rolling drives the verdict: Two projections to keep fresh and a reader who must know which is which; no reset to game and a reportable month alongside
  • RejectedRolling only: Cleanest incentive; leaves the business with no monthly figure and makes compliance reporting a derived estimate
  • RejectedCalendar only: Simplest to explain and to seal; creates the end-of-month cliff and the first-of-month release rush
  • RejectedPublish both and let each service choose which binds: Maximum local autonomy; makes cross-service comparison meaningless and gives every team an argument to have during an incident
Consequences
What it buys
  • No reset date to hold changes against, so the freeze cannot be waited out
  • The business keeps a monthly attainment figure that lines up with every other monthly report
  • An incident's cost visibly decays, which makes the freeze feel like a constraint rather than a punishment
What it costs
  • Two figures for the same SLO, and a reader who confuses them will reach the wrong conclusion
  • A single bad interval can hold a team for four weeks, which is a real organisational cost and the main source of pressure for an override (ADR-11)
  • Twice the materialisation work and twice the recompute on a definition change
Choose differently when
If the organisation's reliability commitments were contractual and monthly — credits owed on a calendar month — the calendar window would have to be canonical, because the figure that carries money cannot be the secondary one. In an internal engineering context the rolling window's incentive properties dominate.
Why it holds up over time
The arithmetic of both windows is permanent and the choice between them is organisational. What should survive is the structural answer: compute both from the same counters, declare one canonical, and name the window in every response. A platform that publishes one figure without saying which window it is will be misread forever.
LessonWhen two framings of the same measurement produce opposite incentives, publish both and declare which one binds. Hiding one does not remove the argument; it just makes it unanswerable.
Shown on views04 12
ADR-15

Admission refuses an unmeasurable objective, an implicit denominator and unbounded cardinality

Accepted

Should the platform accept a definition it knows it cannot measure meaningfully?

Context
Three definitions arrive routinely and all three are worse than useless once published. An objective of 99.999% over a calendar month allows 26 seconds of failure, which minute-grain buckets cannot resolve — the resulting figure is noise presented as precision. A definition that does not say which status classes, synthetic traffic, health checks and client-cancelled requests are excluded has not actually defined anything, and the ambiguity will be resolved later by whoever is under pressure. A definition with unbounded label dimensions is a cardinality incident in the aggregate store, which is the dominant cost driver and the usual cause of a time-series platform becoming unaffordable. In all three cases the cheapest moment to object is before the definition exists.
Decision
Admission runs in the reconciler, before anything enters the registry, and refuses three classes of definition: an objective whose budget is finer than the measurement resolution can resolve, a definition that leaves its valid-event denominator implicit, and label dimensions exceeding a declared cardinality limit. A refusal fails the merge with an explanation; it is not a warning.
How it is realised on AWS
The reconciler computes the implied budget in events and minutes against the window length and bucket resolution, and refuses when the budget falls below the resolvable floor. The schema requires explicit inclusion or exclusion for each status class and traffic type, so an omission is a parse failure rather than a default. label_dims is a bounded list with a per-SLO cell-count cap that also bounds the outcome cube from ADR-03. All three checks run as part of the pull-request validation, so the feedback arrives in review rather than after merge.
Options weighed
  • ChosenRefuse at authoring time, as a merge failure: Friction at the one moment it is cheap; teams will occasionally be blocked from expressing something they wanted
  • RejectedAccept and warn: No friction; the warning is read once and the meaningless figure is published for a year
  • RejectedAccept, and annotate the resulting figures as low-confidence: Honest in principle; nobody reads the annotation and the figure circulates without it
  • RejectedRefuse only the cardinality case: Protects the cost model, which is the self-interested subset; leaves the two that make the figures wrong
Consequences
What it buys
  • No SLO in the estate carries an objective the data cannot support
  • The denominator — where an SLO is really defined and later gamed — is explicit for every SLO by construction
  • Storage cost stays bounded per SLO and attributable to its owner
What it costs
  • A team wanting a very high objective is blocked and may feel the platform is obstructing a legitimate goal
  • The resolution check is arithmetic on assumptions; a change in bucket granularity changes which definitions are admissible
  • Review-time refusal pushes the conversation into a pull request, where the reasoning is less visible than in a design discussion
Choose differently when
If bucket granularity moved from minutes to seconds, the resolvable floor would drop by nearly two orders of magnitude and the first refusal class would mostly disappear — at sixty times the bucket volume. The denominator and cardinality refusals are independent of resolution and would stand regardless.
Why it holds up over time
Refusing to represent what you cannot measure is a permanent property of honest measurement systems. The specific thresholds follow the storage design; the principle that admission is a gate rather than a warning is what stops the registry filling with definitions nobody can defend.
LessonValidate a definition against the limits of your own measurement before you accept it. A system that will happily store an unanswerable question will be asked it, and will answer.
Shown on views02 06 18

Computation and publicationHow the budget is kept fresh at read scale, how a page is decided, and the shape of the document the platform hands over.

ADR-07

Budget state is materialised incrementally behind a watermark, not recomputed per query

Accepted

Can the read path afford the arithmetic the verdict needs?

Context
Computing the budget on every query is the correct-by-construction option: it is always consistent with the current definition, a definition change takes effect instantly, and there is no invalidation logic to get wrong. It is also a scan over up to 40,000 minute buckets per query, and the verdict API is specified at a p99 of 150 ms while absorbing 2,000 requests a second for two minutes during a coordinated release window. Those two facts do not reconcile. Incremental materialisation gives constant-time reads and introduces a staleness window plus invalidation logic that has to stay correct across late data, approved exclusions and retroactive definition changes — three things that all move history.
Decision
Budget state is a materialised projection advanced by a watermark over closed minute buckets. Reads are constant time. The projection carries its watermark, every response carries the resulting computation timestamp, and the verdict API returns insufficient-data once staleness exceeds a declared ceiling rather than serving a figure it cannot date. Invalidation is explicit: late data within the horizon, an approved exclusion and a new definition version each enqueue a scoped recompute (ADR-15, view 15).
How it is realised on AWS
The budget calculator consumes closed buckets from Data Explorer and writes summed counters plus derived percentages into Azure Cache for Redis, keyed by SLO and window kind, with the watermark minute alongside. The verdict API reads Redis only — it never touches the aggregate store on the request path, which is what makes the read tier independently available (ADR-16). Recompute runs in a separate deployment on preemptible capacity so a full-estate backfill cannot consume the capacity serving verdicts.
Options weighed
  • ChosenIncremental materialisation with a watermark and a staleness ceiling: Constant-time reads and an honest staleness figure; invalidation logic that must survive late data, exclusions and version changes
  • RejectedRecompute on read: Always consistent and trivially correct; cannot meet 150 ms p99 at 2,000 rps over a 40,000-bucket scan
  • RejectedRecompute on read with a short-lived response cache: Looks like a compromise; produces the same staleness problem with none of the watermark's visibility into how stale
  • RejectedMaterialise only the rolling window, recompute calendar on read: Halves the projection work; makes the monthly report path a tail-latency risk during exactly the reporting peak
Consequences
What it buys
  • The read path is cheap, flat and independent of history length, so the gate's latency budget is unaffected by retention
  • Staleness is a published number rather than an unknown, and the ceiling turns it into a typed refusal
  • Recompute is a separate concern on separate capacity, which is what makes backfill affordable
What it costs
  • Invalidation is the most intricate logic in the platform, and a missed invalidation serves a stale figure that looks authoritative
  • A definition change is no longer instantaneous: it is a recompute with a visible duration
  • The projection is one more thing that can be wrong in a way the aggregates are not, which is why shadow verification exists (view 15)
Choose differently when
If the verdict read rate were low — a handful of queries a day from a human rather than one per deployment from a machine — recompute-on-read would dominate on correctness and simplicity, and the whole projection could be deleted. The decision is driven entirely by the gate being a machine on the request path.
Why it holds up over time
The specific store will change. What survives is the pairing: if you cache a derived figure, publish its watermark and refuse to serve past a declared staleness ceiling. A cache without a visible age is the mechanism by which correct systems start giving confidently wrong answers.
LessonA derived value served from a cache must carry its own age, and the system must be willing to refuse rather than serve something too old. Freshness you cannot measure is freshness you do not have.
Shown on views08 13 15
ADR-09

Two multi-window burn-rate rules per SLO, generated from a template rather than authored

Accepted

What alerting signal has a defensible relationship to the objective, and how many rules can 1,200 SLOs carry?

Context
Threshold alerting on error percentage is what most organisations have, and it bears no relationship to any commitment: it fires on a bad minute regardless of whether the budget is threatened, and misses a slow leak that quietly exhausts a month. Burn rate fixes the relationship — consumption per unit time, normalised so 1× exhausts the budget exactly at window close — but a single burn-rate rule has to choose between catching a fast outage and catching a slow leak, and cannot do both. The canonical answer is several window and multiplier pairs evaluated simultaneously, which multiplies the rule count across 1,200 SLOs and multiplies the false-positive surface with it. Precision, not recall, is what determines whether pages are still answered in six months.
Decision
Each SLO carries exactly two rules, generated from a template rather than authored per service: a fast pair intended to page, and a slow pair intended to raise a ticket. Both require a short confirmation window to agree with the long window before firing. Page-class and ticket-class alerts route to different destinations, and a slow burn never pages. Every evaluation is recorded with its inputs.
How it is realised on AWS
Rules are compiled from the registry into an alert_rule row per SLO per class — fast at 14.4× over a one-hour long window with a five-minute confirmation window, slow at 3× over six hours — and evaluated by the burn-rate evaluator on separate ticks. Firing decisions go to Azure Monitor action groups, grouped by common dependency with a declared cap on simultaneous pages. alert_decision rows retain the evaluated inputs for 13 months, which is what makes page precision computable against incident records rather than estimated.
Options weighed
  • ChosenTwo generated pairs per SLO, page and ticket, with confirmation: Catches both failure shapes with a bounded rule count; less tuning freedom than per-service authoring
  • RejectedA single fast burn-rate rule: Half the rules and half the noise; misses the slow leak that exhausts a budget over a week, which is the failure nobody notices
  • RejectedThe canonical four-window pairing: Best recall; 4,800 rules across the estate and a false-positive rate that degrades the precision the pages depend on
  • RejectedPer-service authored rules: Maximum local fit; 1,200 hand-written configurations drift immediately and cannot be changed estate-wide
Consequences
What it buys
  • A page means the budget is genuinely going, not that a minute looked bad
  • Alerting behaviour across 400 services can be changed in one merge, which is the only way the configuration stays coherent
  • Page precision is a measured property of the platform rather than an opinion held by the on-call rotation
What it costs
  • Two rules miss failure shapes a four-window pairing would catch, and that gap is accepted deliberately
  • Generated rules fit no service perfectly, so there will be SLOs whose thresholds are visibly wrong for their traffic shape
  • Grouping by common dependency can swallow a genuinely separate incident that happens to share one
Choose differently when
If the on-call rotation were large enough that page volume were not the binding constraint, the four-window pairing's recall would be worth its false positives. At twelve people carrying 1,200 SLOs it is not — and the moment precision drops far enough that pages are routinely ignored, the platform has failed regardless of how many real failures it detected.
Why it holds up over time
Burn rate as the signal is arithmetic and permanent. The rule count is an organisational constraint that will move with the size of the rotation, and the template is the mechanism that lets it move without 1,200 edits. Recording the inputs of every decision is what will still be useful in ten years, when nobody remembers why a threshold was chosen.
LessonAlert on the rate at which a commitment is being consumed, not on the symptom. And measure the precision of your own alerts, because an alerting configuration nobody can evaluate will drift until it is ignored.
Shown on views05 14 17
ADR-12

The platform publishes a signed, typed, time-limited verdict and takes no action

Accepted

Should the reliability authority be able to stop a deployment?

Context
If the platform holds the budget and the policy, the shortest path to a real freeze is for the platform to enforce it: call the deployment API, revoke the pipeline's credential, fail the check itself. It removes an integration and makes the policy unambiguous. It also puts this platform on the critical path of every deployment in the organisation, including the rollback that would end an outage, and it makes a false positive in SLI data an estate-wide inability to ship. The alternative is to publish a claim and let the deployment system act on it, which keeps the authority off the critical path but means the policy is only as real as the gate's willingness to honour it.
Decision
The platform's output is a signed, typed, time-limited verdict document — healthy, warning, exhausted or insufficient-data — carrying the figure, the definition version, the computation timestamp and the data coverage. The platform blocks nothing, revokes nothing and calls no deployment API. Enforcement is the deployment system's act, performed by reading the verdict and applying its own declared policy.
How it is realised on AWS
The verdict API returns a document signed with a non-exportable key in Azure Key Vault Managed HSM, with valid_until set five minutes ahead. The gate authenticates with workload identity federation, verifies the signature itself, and applies the policy recorded against the SLO. Signature validity is short precisely because a signed document is otherwise replayable forever, and the gate's own verification is what makes a cached verdict safe to use during an outage of this platform (ADR-13).
Options weighed
  • ChosenPublish a signed verdict; the gate enforces: One integration and a policy that depends on the gate honouring it; the authority stays off the critical path of the release path
  • RejectedThe platform calls the deployment API to block: Unambiguous enforcement; makes a false positive here an estate-wide inability to ship, including the rollback
  • RejectedThe platform acts as the CI check itself: Neat integration story; the platform's availability becomes the pipeline's availability for every repository
  • RejectedUnsigned verdict over mutual TLS: Simpler and adequate in transit; a cached verdict is then unverifiable, which removes the fail-static option entirely
Consequences
What it buys
  • An outage of the reliability authority is not simultaneously an observability outage and a deployment outage
  • The fail-open or fail-closed decision sits with the system that knows what it is deploying
  • A verdict can be cached, archived, attached to a release record and verified later by anyone
What it costs
  • The policy is only as strong as the gate's implementation, and a bypassable gate makes the arithmetic advisory in fact
  • Signing is on the synchronous read path, so HSM latency and availability are inside the verdict API's own budget
  • Two systems must agree on the verdict schema, and a schema change is a coordinated release
Choose differently when
If deployments were centralised in one pipeline owned by the same team as this platform, direct enforcement would be a smaller risk and would remove a real integration cost. Across 400 services and multiple pipelines it is not, and the rollback argument holds regardless of ownership.
Why it holds up over time
"Publish, do not enforce" is the critical design decision restated at the API boundary, and it will be challenged every few years on the grounds that the integration is tedious and the policy feels weak. The counter-argument does not change with scale: an authority on the critical path of the thing it governs cannot fail independently of it.
LessonMake the output of a judgement a verifiable document rather than an action. Documents can be cached, audited and acted on by systems you do not control; actions cannot, and they put you on everyone's critical path.
Shown on views01 09 13 19
ADR-13

The gate fails static on its cached verdict, and reliability fixes are never frozen

Accepted

What does the release gate do when the platform cannot be reached?

Context
Having decided that the gate enforces (ADR-12), the unreachable case needs an answer, and all three available answers are bad. Fail open evaporates the policy during exactly the incident most likely to take both systems down — a regional event, a shared dependency, the 90-minute recompute window after a failover. Fail closed makes this platform a single point of failure for every deployment in the organisation, including the rollback that would end the outage, which inverts the platform's purpose precisely when it matters. Caching the last verdict and continuing to enforce it turns the problem into a staleness ceiling, which is a judgement with no obviously correct value but at least degrades along a measurable axis.
Decision
The gate caches the last signed verdict it received and, when the platform is unreachable, continues to enforce that verdict up to a declared staleness ceiling, after which it fails open and records that it did so. Separately and unconditionally, changes classed as reliability fixes or rollbacks are exempt from a freeze by default — a freeze that blocks the fix for the outage that caused it is a defect, not a policy.
How it is realised on AWS
The policy object carries gate_mode and an unreachable behaviour, and the verdict's signature plus valid_until let the gate verify a cached document it has held for hours. Rollback and hotfix classes are declared in the policy's exempt_classes and evaluated by the gate from the change's own metadata, not from anything this platform knows. Bypasses and ceiling expiries are reported as first-class figures next to the override register, so a rising bypass rate reads as a policy failure rather than a compliance one.
Options weighed
  • ChosenFail static on a verified cached verdict, with a ceiling and fixes exempt: Degrades along a measurable axis; the ceiling is a judgement nobody can derive, and the policy weakens with age
  • RejectedFail open immediately: Never blocks a fix; removes the policy during the correlated outage that is the likeliest reason the platform is unreachable
  • RejectedFail closed: Strongest policy; makes the reliability platform able to stop the rollback that would end an incident
  • RejectedFail closed with a documented manual override: Looks balanced; puts a manual step into the incident path at the moment humans are least able to execute one
Consequences
What it buys
  • An outage of this platform never prevents a rollback or a reliability fix
  • The policy survives a short outage intact rather than evaporating on the first timeout
  • Degradation is observable: bypasses and ceiling expiries are counted and reviewed
What it costs
  • A gate enforcing an eight-hour-old verdict is enforcing history, and no ceiling value is obviously correct
  • The gate must implement signature verification and a cache, which is real work in every pipeline that integrates
  • Exempting fixes depends on change classification the gate performs, which can be claimed falsely
Choose differently when
If this platform's measured availability were materially better than the deployment system's own, fail closed would stop being a meaningful risk and would give a stronger policy for free. It would still need the rollback exemption, which is the part that is not about availability at all.
Why it holds up over time
Fail-static is the standard answer for a control plane whose consumers must keep working without it, and it will outlive the specific ceiling. The rollback exemption is the part most often forgotten and most expensive to learn during an incident; it should be treated as a permanent invariant rather than a configuration choice.
LessonDecide what a dependent system does without you before you make it depend on you, and never let your own outage block the action that would fix the thing you are measuring.
Shown on views04 13 16 21

Evidence, privilege and recoveryWhat makes the published record trustworthy after the fact, who is allowed to change it, and what happens when the region goes away.

ADR-11

Exclusions annotate rather than delete, and the SLO's author cannot approve one

Accepted

Who is allowed to decide that a bad forty minutes did not count?

Context
Some intervals genuinely should not count: a dependency's declared maintenance, a load test run against production, a measurement artefact already diagnosed. Refusing all exclusions sounds principled and is not — the pressure does not disappear, it relocates, and the next move is to widen the valid-event denominator instead, which is the same adjustment made invisibly. So exclusions have to exist, which creates a privilege whose holder can make any breach appear met. Three mechanisms are available: delete or ignore the interval, annotate it with the original retained, or leave the breach intact and grant a compensating credit. The second preserves the measurement; the third preserves it too but complicates every figure with a second adjustment nobody reads.
Decision
An exclusion is an overlay, never an edit: the interval is annotated with a reason code, a named approver, an expiry and a permanent audit record, and the buckets are untouched so the unexcluded figure stays retrievable forever. The author of an SLO cannot approve an exclusion against it, an exclusion beyond 60 minutes requires a second approver, and exclusions have no permanent form — they expire by default at 30 days.
How it is realised on AWS
exclusion_window rows in Azure SQL carry the interval, reason_code, two approver fields and expires_at; sli_bucket is never modified. The derivation applies the overlay at window reduction, so both figures are a query away. The privileged API refuses an approval whose actor matches the definition's author, writes the audit ledger entry before the registry write, and counts every exclusion into a register reported weekly alongside attainment. Audit rows are Azure SQL ledger tables, so an after-the-fact edit is detectable rather than merely prohibited.
Options weighed
  • ChosenAnnotate, with separation of duties and an expiry: Both figures kept and the adjustment visible; a privilege that still exists and can still be abused by two colluding people
  • RejectedNever exclude anything: Maximally honest in appearance; relocates the pressure into denominator definitions, where the same adjustment is invisible
  • RejectedDelete or ignore the excluded buckets: Simplest to implement and to read; destroys the ability to answer what the month looked like before anyone intervened
  • RejectedBudget credits instead of exclusions: Preserves the measurement perfectly; every figure then carries a second adjustment, and the compound number is one nobody trusts
Consequences
What it buys
  • The question "what did this month look like before the adjustments" always has an answer
  • The privilege to improve a figure never sits with the party being measured
  • A routine exception becomes visible as a rising count rather than as an unremarkable habit
What it costs
  • Two colluding approvers defeat the control entirely, and nothing in the architecture detects collusion
  • Every published figure now has an excluded and an unexcluded form, and a reader who quotes the wrong one is misleading without intending to
  • Exclusion approval is operational work for the reliability lead at exactly the moments they are busiest
Choose differently when
If the published figure ever carried contractual weight — customer credits, a regulatory filing — internal separation of duties would stop being sufficient, because the measured party and the approver are both inside the same organisation. At that point the figure needs an external witness: a notarised snapshot or third-party attestation.
Why it holds up over time
Separation of duties on an adjustment privilege is structural and survives reorganisations in a way that norms about who may adjust a number do not. The specific thresholds will move; the rule that an author cannot approve their own exclusion is the part worth defending.
LessonIf a figure can be adjusted, keep the unadjusted figure forever, put the adjustment behind someone other than the party it benefits, and count the adjustments where the figure is published.
Shown on views15 19 20
ADR-14

At window close an immutable snapshot is sealed and can never be altered

Accepted

If last quarter's attainment is quoted in a board pack, can it change afterwards?

Context
Everything else in this platform is deliberately recomputable: a corrected definition yields a corrected history (ADR-05), an approved exclusion changes the current figure (ADR-11), late data moves a bucket within the horizon (ADR-08). That is correct for an operational figure and wrong for a compliance record. A monthly report regenerated from current state will quietly produce a different answer every time it is run, which makes it useless as evidence — and the circumstances under which it changes are exactly the circumstances under which someone would want it to.
Decision
At the close of each calendar window, a snapshot of the figure — total, consumed, remaining, coverage and the definition version in force — is sealed in write-once storage with no update or delete path. Subsequent definition changes, exclusions and recomputes produce a new current figure and cannot alter a sealed snapshot. The monthly report is rendered from snapshots, never from current state.
How it is realised on AWS
budget_snapshot rows are written once and are immutable in the schema; the rendered snapshot document is written to Azure Blob Storage under a time-based immutability policy with a five-year retention, alongside audit ledger entries in Azure SQL ledger tables. The window manager performs the seal; no other component has write access to that container. A sealed figure that is later found to be wrong is corrected by publishing a dated correction next to it rather than by editing it.
Options weighed
  • ChosenSeal an immutable snapshot at close; correct by addition: A record that means something; a permanent discontinuity in the trend when a long-standing definition error is found
  • RejectedRegenerate reports from current state: Always reflects the best current understanding; a figure that changes between readings cannot be evidence
  • RejectedSnapshot, but allow a privileged correction: Handles the genuinely-wrong case; the correction privilege is exactly the one an interested party would want
  • RejectedSeal with a short grace period before immutability: Pragmatic compromise; the grace period becomes the window in which inconvenient months are revised
Consequences
What it buys
  • A quoted historical figure is stable, dated and attributable to the definition in force at the time
  • An auditor's question is answerable without trusting the platform's own access control
  • The operational figure can stay freely recomputable, because it is no longer doing the compliance job
What it costs
  • A definition error found after a close leaves a permanent discontinuity between the sealed figure and the corrected trend
  • Write-once retention is irreversible by design: a retention policy set wrongly cannot be corrected, with a five-year blast radius
  • Two figures exist for every past month — the sealed one and the current recomputation — and the difference needs explaining
Choose differently when
If the figure were purely internal and never quoted outside engineering, regenerating from current state would be simpler and the loss would be theoretical. The moment a number appears in a board pack or a customer conversation, it needs to be the same number next quarter.
Why it holds up over time
Append-only evidence with correction-by-addition is how accounting has worked for centuries and how every durable audit system works. It will outlive the storage product; the thing to preserve is that the component which seals is the only one with write access, and that nothing has an update path.
LessonSeparate the operational figure from the record. The first should be freely recomputable; the second should be sealed, dated, and corrected only by publishing a correction beside it.
Shown on views11 12 19
ADR-16

Derived state is rebuilt rather than restored, and there is one write region

Accepted

What is actually worth backing up, and can two regions compute the same budget?

Context
The platform holds four kinds of state with very different properties. The registry is irreplaceable and small. The minute buckets are irreplaceable, large and replayable from quarantine. The sealed snapshots and audit ledger are evidence with no update path. The budget projection, compiled rules and report extracts are none of these: they are a deterministic function of the first two, and a backup of them is a slower way to obtain a staler version of something that can be recomputed. Separately, there is a question about regions: a second write region would improve availability and would also have two independent computation planes deriving the same budget from partially overlapping, independently late bucket sets — producing two figures that no key reconciles, in a system whose entire purpose is to produce one.
Decision
Derived state is not backed up at all; its recovery objective is stated as recompute time — a trustworthy projection within 90 minutes of regional recovery. The registry and audit ledger are backed up with a one-minute RPO and geo-replicated; aggregates are replicated and replayable; snapshots are write-once with cross-region redundancy. There is exactly one write region. Failover is promote-then-recompute, and the honest intermediate state is insufficient-data (ADR-04).
How it is realised on AWS
Azure SQL is zone-redundant with point-in-time restore and an active geo-replica in the secondary; the Data Explorer cluster is followed in the secondary; snapshots and audit sit in read-access geo-redundant immutable storage. Redis holds the projection and is not backed up — the recompute runner rebuilds it from aggregates plus registry. The secondary region's compute is scaled to zero. A shadow rebuilder recomputes sampled SLOs continuously and alerts on divergence, because a recovery path exercised only during an incident is a claim rather than a capability.
Options weighed
  • ChosenOne write region; rebuild derived state rather than restore it: One answer always, and a recovery story that improves as compute gets cheaper; 90 minutes of insufficient-data after a regional loss
  • RejectedTwo active write regions: Best availability; two computation planes produce two irreconcilable budgets for the same SLO
  • RejectedOne write region, with the projection backed up for fast restore: Shortens the gap; restores a figure that is stale by exactly the backup interval, which the verdict would have to disclose anyway
  • RejectedActive-passive with synchronous projection replication: Near-zero gap; couples the read path's latency to cross-region replication for state that is disposable by construction
Consequences
What it buys
  • There is never more than one budget figure for an SLO
  • Backup scope is small, so restore drills are cheap and actually get run
  • The recovery story improves automatically as compute gets faster, where a restore of growing data gets slower
What it costs
  • A 90-minute window of insufficient-data after a regional loss, during which every gate falls back to its unreachable behaviour (ADR-13)
  • The secondary's capacity on promotion is a cold-start assumption until a game day proves it
  • Recompute capacity has to be available during a regional incident, which is when it is most contended
Choose differently when
If the verdict API ever became a hard dependency of something that cannot tolerate 90 minutes of insufficient-data — an automated capacity controller, a customer-facing commitment — a second write region would become necessary, and the reconciliation problem would have to be solved rather than avoided, probably by partitioning SLO ownership by region.
Why it holds up over time
Classifying state by rebuildability rather than by importance is the decision that keeps paying off: it shrinks the backup surface, makes recovery testable, and improves with hardware. The one-write-region constraint is specific to needing a single answer and would only change if the arithmetic could be partitioned.
LessonBefore backing something up, ask whether it can be recomputed. If it can, spend the effort on making the recomputation fast and continuously exercised instead of on storing a stale copy.
Shown on views11 15 16
ADR-17

An independent watchdog in a third region holds the platform's own availability signal

Accepted

Who notices when the thing that notices everything stops working?

Context
The platform computes SLOs, so the obvious way to monitor it is to give it SLOs of its own and compute them with the same engine. That is useful and insufficient: the measurement is circular, and the failures that matter most are exactly the ones that take the engine down with the signal. A platform that is up, fast and serving stale or wrong figures passes every internal check, and a platform that is entirely down produces no alert at all because the thing that would produce it is the thing that is down. Nothing inside the system can close that gap, and a self-check deployed inside the same region and the same subscription shares enough failure modes to be nearly as blind.
Decision
The platform holds SLOs on itself — budget freshness, verdict availability, alert detection latency — computed by the same engine for day-to-day operations. Independently, a deliberately minimal watchdog deployed in a third region probes the verdict API from outside and alerts on absence and on staleness, with no dependency on any component it is watching. The package documents that the platform is not the sole authority on its own availability.
How it is realised on AWS
An Azure Functions app in a third region, on its own subscription and its own alerting path, calls the verdict API on a fixed interval with a synthetic read-only SLO, verifies the signature, and checks the computation timestamp against its own clock. It shares no store, no identity and no deployment pipeline with the platform. It raises to the platform rotation through an action group that does not traverse the primary region. Its own availability is explicitly not measured by this platform, which bounds the recursion at one level.
Options weighed
  • ChosenSelf-measured SLOs plus a minimal independent watchdog elsewhere: Bounds the circularity for the cost of a second deployment; the recursion is stopped by fiat rather than solved
  • RejectedSelf-measured SLOs only: No extra component; produces no signal in the one failure mode where a signal is essential
  • RejectedA full second instance of the platform watching the first: Rich signal and genuine independence; doubles the operational surface and raises the same question one level up
  • RejectedRely on the consumers to notice: Zero cost; the gate failing static is silent by design, so nobody notices until a release is questioned
Consequences
What it buys
  • Total unavailability of the platform produces an alert from outside it
  • Staleness is checked against an independent clock, so a stuck computation plane serving old figures is caught
  • The watchdog is small enough to be obviously correct, which is the only useful property for a component of this kind
What it costs
  • A second deployment, subscription and alerting path to maintain, which will be neglected because nothing ever happens to it
  • The watchdog's own availability is unmeasured, so the recursion is stopped by assertion
  • A correlated failure of platform and watchdog is undetected until a human notices
Choose differently when
If an external synthetic monitoring service were already trusted for estate-wide availability checks, it would be the better watchdog — genuinely outside the organisation's cloud estate, and already staffed. Building one is only justified because the probe needs to verify a signature and inspect a computation timestamp rather than just receive a 200.
Why it holds up over time
The need for an external observer of an observer is permanent and the answer is always the same shape: something small, something elsewhere, something with no shared dependencies, and an explicit decision about where the recursion stops. The implementation will be replaced; the reasoning will not.
LessonAny system that reports on others needs something outside it reporting on it, and that something should be small enough that its own correctness is obvious rather than measured.
Shown on views16 17 21

Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on service-level objectives, and the difference matters when reading the decision records.

PackageWhat it isWhat it does hereConsidered instead
Outcome cube The small declared set of counter cells a source emits per SLO per minute — counts by status class and by latency bucket. The emission contract, and the thing that decides which future predicate changes are recomputable (ADR-03). "Metrics", which implies the platform receives arbitrary time series rather than a bounded, declared shape.
Valid-event denominator The explicit statement of which events count toward the SLO at all — which status classes, synthetic traffic, health checks and client-cancelled requests are in or out. Where an SLO is really defined, where it is later quietly widened, and the one thing admission refuses to let a definition leave implicit (ADR-15). "Total requests", which sounds unambiguous and is the source of most disagreement about what an SLO means.
Coverage The fraction of a window's expected minute buckets that were actually present and not no-data. Published alongside every figure, so a number is never read as more certain than the data behind it; below the floor the verdict becomes insufficient-data (ADR-04). "Data quality", which bundles completeness together with correctness and hides which one failed.
no-data A bucket state meaning no valid events were observed for that minute — neither good nor bad, and excluded from both numerator and denominator. The mechanism that stops a stopped metrics pipeline reading as a perfect month. Zero, which an arithmetic pipeline will divide by or treat as success.
insufficient-data A first-class verdict value, returned instead of a figure when coverage is below the floor or the projection is staler than the ceiling. The platform's way of declining to publish something it cannot substantiate; every consumer must handle it as distinct from healthy. An error or a null, both of which consumers routinely treat as "carry on".
Burn rate Error budget consumption per unit time, normalised so that 1× exhausts the window's budget exactly at its close. The only alerting signal with a defensible relationship to the objective (ADR-09). Error rate or error percentage, which fires on a bad minute regardless of whether any commitment is threatened.
Confirmation window A short evaluation window that must agree with the long window before an alert fires. What stops a single anomalous minute paging an on-call engineer. "For duration" or a debounce, which delays an alert without requiring a second independent agreement.
Verdict A signed, typed, time-limited document stating healthy, warning, exhausted or insufficient-data, with the figure, definition version, computation timestamp and coverage. The platform's entire output to machines. It is a claim, not an instruction (ADR-12). "Gate result" or "check status", which imply the platform performed the enforcement.
Watermark The most recent minute the budget projection has incorporated. What makes staleness measurable, and therefore what makes the staleness ceiling and the insufficient-data refusal possible (ADR-07). "Last updated", which is usually the time of the write rather than the time of the data.
Exclusion window An approved interval annotated as outside the valid-event denominator, with a reason code, two approvers and an expiry. The sanctioned way to say a bad interval should not count, designed so the unexcluded figure survives (ADR-11). "Maintenance window" or "downtime exemption", which imply the data was removed rather than annotated.
Sealed snapshot The immutable record of a calendar window's figure, written once at close with no update path. The compliance record, as distinct from the freely recomputable operational figure (ADR-14). "Monthly report", which in most systems is regenerated from current state and therefore changes between readings.
Fail static The release gate continuing to enforce its last verified cached verdict while the platform is unreachable, up to a declared staleness ceiling. The declared answer to the unreachable case, chosen over fail-open and fail-closed (ADR-13). "Fail safe", which is ambiguous here — open and closed are each the safe option under a different failure.
Derived state The budget projection, compiled alert rules and report extracts — a deterministic function of the registry and the minute buckets. Deliberately not backed up; its recovery objective is recompute time rather than an RPO (ADR-16). "Cache", which understates it — consumers read it as authoritative, so its staleness has to be published.
Measurement failure An alert class meaning the platform cannot see an SLO, as distinct from an alert meaning the service is failing. Routed to the platform rotation rather than the service rotation, so triage starts at the broken system (ADR-10). A low-severity reliability alert, which still points the reader at the wrong system.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.