# SLO and Error Budget Service — Architecture One-Pager and Decision Record

*SLO and Error Budget Service · Solution Architecture v1.0 · Microsoft Azure · Reliability Architecture · 2026-10 · 21 views · 17 architecture decision records*

The argument these decisions serve is summarised in the [Architecture One-Pager](architecture-one-pager).

The architecture behind the Monday-morning argument. The status page is green, the dashboards are green, and a dozen weekend tickets say checkout failed — and someone has to decide whether the release train ships. This platform replaces conviction with arithmetic: for each critical user journey it holds a declared reliability target, measures what users actually got, and publishes the difference as an error budget with a typed verdict attached. It is built on Microsoft Azure, and it is deliberately the thinnest useful thing in the estate: it collects no telemetry and it stops no deployments.

> **Status of this document.** This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure in this package is a stated assumption, chosen to be defensible and arguable rather than measured. They are stated precisely so that a reviewer can disagree with one and follow it to the decision that depends on it. The requirement document marks them as assumptions section by section; the operating context assumed throughout is 1,200 SLOs across 400 services and 60 critical user journeys, growing 30% a year, with 1.73 million minute buckets written a day and a verdict API serving 400 requests a second steady state and 2,000 for two minutes during a coordinated release window.

## How to read a record

- **Question:** The forcing question: why a decision was needed at all.
- **Context:** The requirement, the scale and the constraint that make it hard.
- **Decision:** What this architecture does, stated so it can be checked.
- **How it is realised on AWS:** The concrete mechanism: which service or package, configured how, in which subscription.
- **Options weighed:** Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- **Consequences:** What the choice buys and what it costs, both kept visible.
- **Choose differently when:** The conditions that would flip the decision for your system.
- **Why it holds up over time:** What keeps the decision right as scale, staff and technology change.
- **Lesson:** The principle that transfers beyond this platform.

## Decision map

**The derivation boundary**: The decision that defines the architecture, and the two that follow from it directly: what the platform is allowed to own, and what it is allowed to store.

- ADR-01 · Good and valid counts are stored separately at every level; no ratio is ever stored
- ADR-02 · The platform consumes aggregates and never collects, samples or stores telemetry
- ADR-03 · Sources emit a low-cardinality outcome cube, not a single good/valid pair

**Missing data as a first-class answer**: What the platform says when it cannot substantiate a figure, which is the difference between a reliability authority and a reassurance machine.

- ADR-04 · An empty bucket is no-data; below the coverage floor the verdict is insufficient-data
- ADR-08 · Buckets are assigned by event time with a 10-minute horizon, and idempotency is in the primary key
- ADR-10 · Measurement failure is a separate alert class, routed to a different rotation

**Definition, admission and windows**: Where an SLO comes from, what the platform refuses to accept, and which window the freeze is computed over.

- ADR-05 · SLO definitions are versioned immutable records reconciled from the owning team's repository
- ADR-06 · Rolling and calendar windows are both published; the rolling window drives the verdict
- ADR-15 · Admission refuses an unmeasurable objective, an implicit denominator and unbounded cardinality

**Computation and publication**: How the budget is kept fresh at read scale, how a page is decided, and the shape of the document the platform hands over.

- ADR-07 · Budget state is materialised incrementally behind a watermark, not recomputed per query
- ADR-09 · Two multi-window burn-rate rules per SLO, generated from a template rather than authored
- ADR-12 · The platform publishes a signed, typed, time-limited verdict and takes no action
- ADR-13 · The gate fails static on its cached verdict, and reliability fixes are never frozen

**Evidence, privilege and recovery**: What makes the published record trustworthy after the fact, who is allowed to change it, and what happens when the region goes away.

- ADR-11 · Exclusions annotate rather than delete, and the SLO's author cannot approve one
- ADR-14 · At window close an immutable snapshot is sealed and can never be altered
- ADR-16 · Derived state is rebuilt rather than restored, and there is one write region
- ADR-17 · An independent watchdog in a third region holds the platform's own availability signal

## Technology by capability

Microsoft Azure was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's reliability-and-operations use cases the observability platform and health-check service are on AWS, chaos engineering is on Google Cloud, and incident management and backup are self-hosted open source — so this area had no Azure document, and reaching for the same cloud repeatedly teaches a service catalogue rather than an architecture. The second is that Azure has no first-class managed SLO and error-budget product the way Google Cloud's monitoring does, which makes it the more instructive substrate: the burn-rate engine, the window manager and the verdict signing all have to be designed rather than configured, and the decisions stay visible. Every requirement in ask.md is vendor-neutral; the table below is one defensible realisation of it, and the capability column is what a reviewer should argue with.

| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| SLO registry, policies, exclusions, overrides | Azure SQL Database, zone-redundant, point-in-time restore, active geo-replica, ledger tables for audit | Microsoft Azure | PostgreSQL Flexible Server with pgaudit; Cosmos DB with strong consistency; CockroachDB self-hosted | The only state needing transactions, foreign keys and a tamper-evident audit path in one engine; small enough that its durability is cheap | ADR-05 |
| SLI minute-bucket aggregates | Azure Data Explorer, time-partitioned, hot/warm tiering, 13-month retention | Microsoft Azure | ClickHouse self-hosted; Mimir or Thanos; Synapse dedicated pools; TimescaleDB | Columnar range scans over a 28-day window per SLO are the dominant read pattern, and update policies plus materialised views cover the derivation | ADR-01 |
| Indicator stream | Azure Event Hubs, partitioned by SLO identity, with consumer checkpointing | Microsoft Azure | Kafka or Confluent Cloud; Service Bus topics; Kinesis | Per-key ordering and replayable offsets, which is what makes the idempotent bucket upsert in ADR-08 safe after a consumer crash | ADR-08 |
| Budget projection | Azure Cache for Redis, zone-redundant, no backup | Microsoft Azure | Cosmos DB single-region; in-process cache with a rebuild on start; Data Explorer materialised view read directly | Constant-time reads at the verdict API's burst rate; deliberately disposable, with recovery expressed as recompute time (ADR-16) | ADR-07 |
| Computation, ingest, publication and control tiers | Azure Container Apps with KEDA scaling, separate apps per tier | Microsoft Azure | AKS; App Service; Azure Functions; Kubernetes anywhere | Independent deploy and scale per tier without operating a cluster — a reliability platform should not be the most operationally demanding thing in the estate | ADR-07 |
| Recompute and backfill | Container Apps jobs on spot-priced capacity, scheduled | Microsoft Azure | Batch; AKS spot node pools; Databricks jobs | A full-estate rebuild is interruptible and off the read path, so preemptible capacity is the right price for it | ADR-16 |
| Window snapshots and audit evidence | Blob Storage with time-based immutability, RA-GRS, five-year retention | Microsoft Azure | S3 Object Lock; GCS bucket lock; WORM appliance | Write-once with no delete path is the whole requirement in ADR-14; cross-region read access makes the evidence survive the write region | ADR-14 |
| Verdict signing | Azure Key Vault Managed HSM, non-exportable signing key | Microsoft Azure | Key Vault standard tier; CloudHSM; an offline key with short-lived intermediates | A verdict a machine acts on must be verifiable by a party that does not trust this platform's network path (ADR-12) | ADR-12 |
| Identity for humans and machines | Microsoft Entra ID with SSO, MFA, and workload identity federation for the release gate | Microsoft Azure | Okta with OIDC federation; SPIFFE/SPIRE for workloads | Four independently assignable privileges with no long-lived API keys, and OIDC federation so the gate holds no secret | ADR-11 |
| Edge, quota and throttling for the read path | Azure Front Door plus API Management | Microsoft Azure | Application Gateway with a WAF; NGINX ingress; Cloudflare | The verdict API absorbs a 5× burst during release windows and needs per-caller quota without that logic living in the service | ADR-13 |
| Upstream measurement plane | Azure Monitor and Azure Monitor managed service for Prometheus, with recording rules emitting the outcome cube | Microsoft Azure | Self-hosted Prometheus with remote write; Datadog; Grafana Cloud | Deliberately outside the platform (ADR-02); recording rules are where the outcome cube in ADR-03 is produced | ADR-02 |
| Alert routing | Azure Monitor action groups, separate groups per alert class | Microsoft Azure | PagerDuty or Opsgenie directly; Alertmanager | Two alert classes must reach two rotations by separate paths (ADR-10), and the measurement-failure path must not traverse the primary region | ADR-10 |
| Independent watchdog | Azure Functions in a third region, separate subscription and alerting path | Microsoft Azure | An external synthetic monitoring service; a probe in another cloud; Grafana Synthetic Monitoring | Small enough to be obviously correct, and sharing no store, identity or pipeline with what it watches (ADR-17) | ADR-17 |
| Reporting | Scheduled extract from sealed snapshots to a reporting sink | Microsoft Azure | Live Power BI against Data Explorer; a static site generated per month | Reports are rendered from immutable snapshots rather than current state, so a quoted figure does not move between readings (ADR-14) | ADR-14 |

## The decisions, and the alternatives that lost

### The derivation boundary

*The decision that defines the architecture, and the two that follow from it directly: what the platform is allowed to own, and what it is allowed to store.*

#### ADR-01 · Good and valid counts are stored separately at every level; no ratio is ever stored

**Status:** Accepted  ·  **Shown on views:** 10, 12

*If someone asks for last quarter's availability, can the platform compute it from what it kept?*

**Context.** The natural thing to store is the answer: a availability percentage per minute, per hour, per day. It is what people want to read, it is what every dashboard displays, and it is one number instead of two. It is also unusable. A ratio cannot be re-aggregated: the mean of twelve hourly availability figures is not the day's availability unless every hour carried identical traffic, and traffic is never identical. Worse, a stored ratio has thrown away the denominator, so there is no way to express a budget in the only unit that is actually spendable — failed requests. Every later capability in this design depends on being able to re-reduce the same underlying counts over a different window, a different exclusion overlay or a different predicate, and all of them become impossible rather than merely expensive the moment the division happens before storage.

**Decision.** The unit of storage is a pair of integers per SLO per minute: good_count and valid_count, plus a coverage flag. Ratios are computed at read time, never persisted. Every window figure — rolling, calendar, attributed interval, excluded or unexcluded — is a sum of the same two columns over a different set of rows.

**How it is realised on AWS.** The sli_bucket table in Azure Data Explorer is keyed on (slo_id, minute_utc) with good_count, valid_count, coverage and source_class. Window reduction is a Kusto summarise over a time range; the budget is total minus consumed where both are counts. The budget projection in Redis holds the summed counters and the derived percentage together, with the counters authoritative — a reader that wants to recombine windows reads the counters, and a reader that wants to display a figure reads the percentage.

| Option | Verdict | Reasoning |
|---|---|---|
| Store good and valid counts separately | Chosen | Two integers per bucket instead of one float; everything downstream stays recomputable and the budget is expressible in failed requests |
| Store the ratio per bucket | Rejected | Halves the storage and makes correct re-aggregation across windows impossible, which silently removes rolling windows, exclusion overlays and interval attribution |
| Store raw request events and aggregate at query time | Rejected | Maximum flexibility at two orders of magnitude more storage and ingest; the platform would then own collection, contradicting ADR-02 |
| Store counters plus a pre-computed ratio as the authority | Rejected | Invites readers to trust the ratio and quietly reintroduces the re-aggregation bug the counters were kept to avoid |

**What it buys**

- Any window, any exclusion overlay and any corrected predicate can be recomputed from retained data
- The budget is expressible as a count of permitted failures, which is what makes it feel spendable rather than abstract
- Interval attribution is a sort over the same two columns rather than a separate pipeline

**What it costs**

- Two columns to keep consistent, and a reader who sums ratios by hand will still get a wrong answer the platform cannot prevent
- The percentage everyone wants to read is computed on every query, so the read path carries arithmetic a stored ratio would have avoided

**Choose differently when.** If the platform only ever had to answer "what is the availability of this window, right now" — no rolling windows, no exclusions, no retroactive correction, no interval attribution — a stored ratio would be adequate and cheaper. In practice the first request after launch is always "what happened on Saturday", which needs the counts.

**Why it holds up over time.** This is arithmetic, not engineering. Any future store can re-aggregate counters and none can un-divide a ratio, so this decision should outlive every technology named in the package. It is also the decision most often skipped, because the ratio is the thing people want to see and the cost of storing it looks like a saving.

> **Lesson.** In any system that aggregates, store the operands and not the result. The result is cheap to recompute and impossible to decompose, and the capabilities you will want in six months are all decompositions.

#### ADR-02 · The platform consumes aggregates and never collects, samples or stores telemetry

**Status:** Accepted  ·  **Shown on views:** 01, 07, 21

*When the reliability authority and the measurement pipeline fail together, what is left?*

**Context.** The commercially obvious design has the SLO platform own collection: agents in the fleet, its own time-series store, its own query language, SLOs computed over data it controls end to end. It is what several products do, it removes an integration, and it tightens the latency from event to figure. It also makes one system simultaneously the measurement path and the reliability authority, which has two consequences. Operationally, its outage is both an observability outage and a loss of the signal that would tell you an observability outage is happening. Epistemically, it can no longer distinguish "I did not receive data" from "there was no bad traffic", because it is the thing that would have received it — and a platform that cannot represent its own blindness will report blindness as health.

**Decision.** Collection belongs to the observability platform. This platform's inputs are pre-aggregated counters arriving on a stream, and its outputs are derived figures and signed verdicts. It retains aggregates and derivations; it retains no raw signal, no traces and no logs of the measured systems. Enforcement, symmetrically, belongs to the deployment system (ADR-12).

**How it is realised on AWS.** Azure Monitor and Azure Monitor managed service for Prometheus remain the system of record for telemetry. Recording rules in the measurement plane emit the outcome cube to Event Hubs; the platform's bucketer consumes it. Nothing in the platform scrapes, and no agent is deployed to a measured workload. The layered view draws the measurement plane as the bottom band, labelled "outside", so no reader of the set mistakes it for a component.

| Option | Verdict | Reasoning |
|---|---|---|
| Consume aggregates; own no collection | Chosen | One integration and a dependency on someone else's uptime, in exchange for reproducibility and the ability to describe blindness |
| Own collection end to end | Rejected | Tighter latency and no integration; makes the reliability authority a hard dependency of the measurement path and unable to represent its own blindness |
| Consume aggregates but keep a parallel collection path for critical journeys | Rejected | Two sources of truth for the same SLO, and the first disagreement between them is unresolvable |
| Buy a managed SLO product that owns both collection and enforcement | Rejected | Fastest to a figure; concentrates measurement, derivation and enforcement in one vendor outage |

**What it buys**

- A measurement outage is describable as a coverage deficit rather than invisible as a good month
- Every published figure is a reproducible function of retained aggregates plus a versioned definition
- The platform stays small enough that it is not the most operationally demanding thing in the estate

**What it costs**

- The platform cannot fix bad telemetry; it can only refuse to launder it, which moves some problems to a team that does not own the budget
- The outcome cube's dimensions are chosen in a system this platform does not control (ADR-03)
- Event-to-figure latency carries an extra hop, which is why budget freshness is specified at 60 s rather than seconds

**Choose differently when.** If blind minutes became routine rather than exceptional — coverage regularly below the floor across many SLOs — the boundary would stop paying for itself, because a platform that answers insufficient-data most of the time is not an authority. At that point owning collection for the critical journeys would be the correct response, and the decision should be formally revisited rather than quietly eroded.

**Why it holds up over time.** "Compute and publish, do not measure and do not enforce" is a statement about which systems may depend on which, not about any product. It is also the single most useful test to apply to future proposals: a vendor platform that owns collection and gating contradicts it, and that contradiction is the thing to argue about rather than the feature list.

> **Lesson.** Do not let the system that judges a thing also be the system that observes it and the system that punishes it. Separating the three is what makes each one's failure describable by the others.

#### ADR-03 · Sources emit a low-cardinality outcome cube, not a single good/valid pair

**Status:** Accepted  ·  **Shown on views:** 06, 10

*When a team wants to change what "good" means, how much history can they take with them?*

**Context.** Having decided to consume aggregates (ADR-02), the obvious contract is the one the platform needs: each source emits good_count and valid_count per minute per SLO. It is minimal and cheap. It also freezes the good-event predicate at the point of emission, which means the first time a team wants to tighten a latency threshold from 800 ms to 500 ms, there is no history to recompute against — the distinction was discarded upstream before the platform ever saw it. Request-level events would preserve everything, at two orders of magnitude more ingest and storage, and would drag the platform back across the collection boundary. The middle position is to notice that almost every predicate change teams actually want moves along one of two axes: which status classes count as failures, and which latency threshold counts as slow.

**Decision.** Sources emit a small declared cube per SLO per minute — counts bucketed by status class and by latency bucket — rather than a single good/valid pair. The good-event predicate becomes a projection over the cube's declared dimensions, so a predicate change along those dimensions is recomputable over retained history at pre-aggregation cost. A change outside the declared dimensions is not recomputable, and the platform says so rather than pretending.

**How it is realised on AWS.** Prometheus recording rules in the measurement plane emit a histogram-shaped record: for each SLO and minute, counts keyed by status class (2xx, 4xx, 5xx, client-cancelled) and by latency bucket on a fixed ladder. The bucketer flattens the cube into sli_bucket rows keyed by (slo_id, minute_utc) with the cube retained as a nested column, so window reduction sums the cells the current predicate selects. Cardinality is bounded by the admission check in ADR-15: the cube's cell count per SLO per minute is capped, which is what keeps this affordable at 1.73 million buckets a day.

| Option | Verdict | Reasoning |
|---|---|---|
| A small declared outcome cube | Chosen | A few cells per bucket instead of two integers; keeps retroactive redefinition possible along the axes teams actually use |
| A single good/valid pair | Rejected | Cheapest contract; makes every predicate change forward-only, which is the complaint that arrives in month two |
| Request-level events into the platform | Rejected | Any predicate, any time, full attributability; two orders of magnitude more cost and it crosses the collection boundary (ADR-02) |
| Cube for critical journeys, pair for the long tail | Rejected | Tempting on cost; produces two classes of SLO with different capabilities and a migration whenever a journey is promoted |

**What it buys**

- Tightening a latency threshold or reclassifying a status code recomputes over 13 months of history with no upstream change
- Cost stays within a small multiple of the minimal contract, because the cube is bounded rather than open
- The limits of recomputability are explicit — a team can be told which changes are free and which need a source change

**What it costs**

- The cube's dimensions are chosen once, upstream, by a team that does not own this platform, and a dimension nobody thought to emit is unrecoverable
- A nested column per bucket is more complex to reduce than two integers, and the reduction has to stay deterministic for ADR-01 to hold
- Teams will occasionally want a predicate the cube cannot express, and the answer is a genuine upstream change rather than a configuration one

**Choose differently when.** If storage and ingest for request-level events became cheap enough that 2.5 million events a second could be retained for 13 months without a separate budget conversation, the cube would be an unnecessary compromise and event-level ingest would dominate it — though the collection-boundary objection in ADR-02 would still stand, so the events would have to arrive from the measurement plane rather than be collected here.

**Why it holds up over time.** The cube is a bet about which questions get asked, not about a technology. The specific dimensions will be revised; the principle — pre-aggregate along the axes along which the definition is likely to move, and declare the limits — transfers to any pre-aggregation decision in any system.

> **Lesson.** When you pre-aggregate, you are choosing which future questions remain answerable. Spend the extra dimension on the axis your definitions will move along, and say plainly which questions you have made unanswerable.

### Missing data as a first-class answer

*What the platform says when it cannot substantiate a figure, which is the difference between a reliability authority and a reassurance machine.*

#### ADR-04 · An empty bucket is no-data; below the coverage floor the verdict is insufficient-data

**Status:** Accepted  ·  **Shown on views:** 05, 10, 17, 21

*What does the platform publish when it did not receive the data?*

**Context.** A minute with zero valid events is ambiguous: it may be three in the morning on a low-traffic journey, or the metrics pipeline may have stopped. The convenient implementations are both wrong. Treating the minute as good — a ratio of 0/0 rendered as 100% — means a stopped scrape presents as a perfect month, which is the single most common way this class of platform lies. Treating it as bad means every quiet night burns budget and the SLO becomes unusable for anything but the busiest journeys. Interpolating from neighbouring minutes is worse than either, because it manufactures evidence. The deeper point is that a reliability authority has an asymmetric cost function: a wrong figure is materially worse than no figure, because a wrong figure is acted on.

**Decision.** A bucket with no valid events is recorded in a third state, no-data, which is neither good nor bad and is excluded from both numerator and denominator. Coverage — the fraction of expected buckets present — is computed per window and published alongside every figure. Below a declared coverage floor the verdict is the typed value insufficient-data rather than a number, and reliability alerts for that SLO are suppressed in favour of a measurement-failure alert (ADR-10).

**How it is realised on AWS.** sli_bucket carries a coverage column with values ok and no-data. Window reduction counts expected minutes against present-and-ok minutes to produce coverage_pct, which is written into the budget projection and returned in every verdict and query response. The verdict API evaluates coverage before anything else and short-circuits to insufficient-data below 95%. Low-traffic journeys declare an aggregation rule — a minimum valid-event count per evaluation period — rather than being silently imputed.

| Option | Verdict | Reasoning |
|---|---|---|
| no-data as a third state, with a published coverage floor | Chosen | Costs a column, a floor nobody can derive from first principles, and a fourth verdict value; makes blindness representable |
| Treat an empty bucket as good (0/0 = 100%) | Rejected | Simplest and the most dangerous: a stopped pipeline is indistinguishable from a perfect month |
| Treat an empty bucket as bad | Rejected | Fails safe in principle; makes every low-traffic journey permanently in breach and the platform unusable for the long tail |
| Interpolate from neighbouring minutes | Rejected | Produces a plausible figure with no evidence behind it, which is the one outcome worse than no figure |

**What it buys**

- A metrics outage is visible as itself rather than as good news
- Every figure carries its own confidence, so "11% remaining" is never read as more certain than the data behind it
- Regional recovery has an honest intermediate state: insufficient-data for 90 minutes while the projection rebuilds (ADR-16)

**What it costs**

- The 95% floor is a judgement with no principled derivation, and SLOs will sit just above and just below it for reasons unrelated to reliability
- insufficient-data is a fourth state every consumer has to handle, and a consumer that treats it as healthy has reintroduced the original bug
- A low-traffic journey needs an explicit aggregation rule, which is extra work at definition time

**Choose differently when.** If every SLO in the estate were on a journey with continuous high traffic and a measurement pipeline with its own hard availability guarantee, the ambiguity would largely disappear and a two-state bucket would be adequate. That is not true of the long tail of 400 services, and it is never true during an incident — which is exactly when the figure is read.

**Why it holds up over time.** Telemetry pipelines will keep breaking in new ways indefinitely. A design in which absence is representable degrades gracefully under failures nobody has imagined yet; one in which absence is indistinguishable from success fails silently under all of them. This should be among the last decisions anyone revisits.

> **Lesson.** Give every measurement system an explicit way to say "I do not know", and make the consumers handle it. A system that can only report good or bad will report its own blindness as good.

#### ADR-08 · Buckets are assigned by event time with a 10-minute horizon, and idempotency is in the primary key

**Status:** Accepted  ·  **Shown on views:** 10, 12, 15

*What happens when the same minute's data arrives twice, and when it arrives too late?*

**Context.** Aggregates arrive over a stream, which means they arrive at-least-once, out of order, and sometimes long after the minute they describe. Assigning them to buckets by arrival time is simpler and wrong: a backlog drained after a pipeline incident would attribute an hour of old failures to the minute the backlog cleared, which both hides the real interval and manufactures a fake one. Event-time assignment is correct but leaves two questions: how long a bucket stays open for late arrivals, and what stops a replayed batch from double-counting. An unbounded reopen window means no figure is ever final; an exactly-once delivery guarantee is not available from any stream the platform would use.

**Decision.** Every aggregate is assigned to its bucket by event time. A bucket accepts late arrivals for 10 minutes after its close and is then immutable; data arriving later is recorded as a coverage deficit rather than silently dropped or silently applied. Idempotency is structural: (slo_id, minute_utc) is the bucket's primary key, so a replayed batch is a key collision resolved by overwrite-with-identical-value rather than an addition.

**How it is realised on AWS.** The bucketer consumes Event Hubs partitions with checkpointing, derives minute_utc from the record's own timestamp, and upserts sli_bucket rows on (slo_id, minute_utc) — the write is the full cell set for that minute from that source, not a delta, which is what makes a replay a no-op. A late record whose bucket has passed the horizon increments the window's coverage deficit. An input batch that would move a closed window's figure beyond a declared tolerance is quarantined with a typed reason rather than applied, and the quarantine is a retained, replayable store rather than a dead-letter queue.

| Option | Verdict | Reasoning |
|---|---|---|
| Event time, 10-minute horizon, idempotent upsert on (slo_id, minute) | Chosen | Correct attribution, bounded finality, and double-counting made impossible by the key rather than by the consumer's care |
| Arrival-time bucketing | Rejected | Trivial to implement; a drained backlog invents an incident and hides the real one |
| Event time with an unbounded reopen window | Rejected | Never loses a record; no figure is ever final, so nothing can be sealed and the compliance record becomes provisional forever |
| Additive deltas with a de-duplication table | Rejected | Standard streaming pattern; replaces a primary key with a second system that has to be correct, and its failure mode is a silently wrong count |

**What it buys**

- An interval attribution points at the minutes users actually suffered, which is what makes view 04's trough answerable
- A replay — after a bucketer crash, a source backfill or an operator retry — cannot change a figure
- Windows become final, so snapshots can be sealed and the compliance record means something (ADR-14)

**What it costs**

- Data later than 10 minutes is lost as measurement and visible only as reduced coverage, which makes the figure slightly pessimistic
- The upsert must carry the full cell set per source per minute, so the write is larger than a delta
- Multiple sources writing the same SLO's minute need source_class in the key path or they overwrite each other

**Choose differently when.** If the upstream pipeline offered a hard bound on delivery latency — every record within 30 seconds or not at all — the horizon could shrink to near zero and the coverage-deficit path would be mostly dead code. Conversely, if the platform ever had to accept bulk historical imports as a normal operation, the horizon would have to become a per-source declaration rather than a constant.

**Why it holds up over time.** Event-time semantics and key-based idempotency are the two things stream processing has converged on for good reasons, and neither depends on a product. The horizon's specific length is a tuning parameter that will change; the existence of a declared horizon, after which a figure is final, is what makes everything downstream possible.

> **Lesson.** Put idempotency in the primary key, not in the consumer's logic. A duplicate that collides is a non-event; a duplicate that has to be detected is a bug waiting for load.

#### ADR-10 · Measurement failure is a separate alert class, routed to a different rotation

**Status:** Accepted  ·  **Shown on views:** 05, 14, 17, 21

*When the SLO goes blind, who is woken, and what are they told?*

**Context.** Once blindness is representable (ADR-04), the alerting question follows: what happens when coverage collapses? Doing nothing is unacceptable — a silently blind SLO is an unmonitored journey. Firing the reliability alert is worse than nothing, because it tells the on-call engineer that users are suffering when the truth is that the platform cannot see. That is the single worst minute in the whole system: the engineer is awake, under time pressure, and holding a signal that points at the wrong system entirely. The diagnosis they need — "this is a metrics problem, not a service problem" — is a different conclusion, reached by a different person, with a different runbook.

**Decision.** Below the declared coverage floor, reliability alerts for that SLO are suppressed and a distinct measurement-failure alert is raised instead, routed to the platform rotation rather than the service rotation. The two classes are never merged, and the measurement-failure alert names the SLO, the coverage figure and the source class that stopped.

**How it is realised on AWS.** The burn-rate evaluator checks coverage before evaluating any rule and emits one of two alert types. Measurement-failure alerts route through a separate Azure Monitor action group to the platform rotation; reliability alerts route to the owning service's rotation. The suppression is recorded on the alert_decision row with suppressed_reason, so a suppressed window is explicable afterwards and does not silently vanish. Nights attributed to measurement failure are excluded from the page-precision figure in ADR-09, so a broken pipeline cannot degrade the statistics that justify the alerting configuration.

| Option | Verdict | Reasoning |
|---|---|---|
| Two alert classes, separate rotations, with suppression recorded | Chosen | An extra alert type and routing to maintain; the engineer is told which system is broken |
| One alert class covering both | Rejected | Simplest routing; puts every reader in the position of being unable to tell an outage from a blind SLO |
| Suppress silently below the floor | Rejected | No false pages at all; an unmonitored journey nobody is told about, which is the failure this platform exists to prevent |
| Degrade the reliability alert's severity instead of reclassifying | Rejected | One mechanism, less plumbing; a low-severity reliability alert still points the reader at the wrong system |

**What it buys**

- The page names the broken system, so triage starts in the right place
- Pipeline health becomes a platform concern with its own rotation rather than an unexplained service symptom
- Alert precision stays a meaningful figure, because measurement failures do not pollute it

**What it costs**

- Two rotations must both exist and both be staffed, which is an organisational dependency this design assumes
- An SLO can sit blind and un-alerted to its owning team, who find out from the coverage figure rather than from a page
- A partial measurement failure above the floor still produces a reliability alert computed on thin data

**Choose differently when.** If one team owned both the measurement pipeline and the services being measured, the split would be bureaucratic rather than useful and a single rotation with a well-worded alert body would be adequate. The split exists because the two systems have different owners.

**Why it holds up over time.** The distinction between "the thing is broken" and "my view of the thing is broken" is permanent in any monitoring system. Products and routing will change; the requirement that those two conclusions reach different people with different runbooks will not.

> **Lesson.** Never let a monitoring system report its own blindness through the same channel it reports failures. The reader cannot distinguish them, and they will act on the wrong one under time pressure.

### Definition, admission and windows

*Where an SLO comes from, what the platform refuses to accept, and which window the freeze is computed over.*

#### ADR-05 · SLO definitions are versioned immutable records reconciled from the owning team's repository

**Status:** Accepted  ·  **Shown on views:** 06, 12, 18

*When a target turns out to have been defined wrongly for three months, what happens to those three months?*

**Context.** A console where SLOs are edited in place is the fastest thing to build and the easiest to use. It also makes a definition change indistinguishable from a definition correction, destroys the association between a published figure and the text that produced it, and leaves no review step at the one moment when review matters most — because an SLO is a commitment, and widening a denominator is a commitment change dressed as a configuration tweak. The alternative is to treat the definition as code: text in the owning team's repository, reviewed by pull request, reconciled into a registry that is a projection of it rather than a system of authorship.

**Decision.** SLO definitions are declarative text in the owning team's repository. A reconciler reads merged definitions and writes versioned, immutable records into the registry, each with an effective-from timestamp. Prior versions remain readable forever, so every historical figure stays attributable to the exact text that produced it, and a correction produces a corrected history rather than a silent rewrite.

**How it is realised on AWS.** A Container Apps job polls the definitions repository and reconciles merged changes into Azure SQL, writing a new slo_version row rather than updating one. slo carries current_version as an FK; every budget_window and verdict records the version_id it was computed under. The admission checks in ADR-15 run in the reconciler, so a definition that cannot be measured fails the merge rather than entering the registry. The console is read-mostly: it can pause an SLO and request an exclusion, but it cannot author a predicate.

| Option | Verdict | Reasoning |
|---|---|---|
| Repository-authored text, reconciled into a versioned registry | Chosen | Review at the right moment, permanent attribution, and corrections that are visible as corrections; slower to change than a console |
| Console editing with an audit log | Rejected | Fast and comfortable; the audit log records that a change happened without making the prior definition's figures recomputable |
| Versioned records authored through an API, no repository | Rejected | Keeps immutability and loses the review step, which is most of the value — nobody reviews an API call |
| Repository as the only store, no registry | Rejected | Purest GitOps; makes every verdict read a repository read, and the registry's transactional guarantees are needed for exclusions and policy |

**What it buys**

- A wrong definition is corrected and the history recomputes, instead of the error being frozen into the record
- Widening a denominator is a reviewed change with a named author, which is the only structural counterweight to targets drifting easier
- Every figure in a monthly report can be traced to the text in force at the time

**What it costs**

- Changing an SLO takes a pull request, which teams will find slow in exactly the moment they want it to be fast
- The registry can lag the repository, so there is a reconciliation window during which the two disagree
- Definition ownership stays with the service team, which makes targets achievable and comparability across 400 services weak

**Choose differently when.** If SLO definitions needed to change inside minutes — an incident-time adjustment, a dynamically generated SLO per tenant — a repository round trip would be the wrong mechanism and an API with versioned records would dominate. Nothing in this problem has that shape: an SLO that changes hourly is not a commitment.

**Why it holds up over time.** "The definition is reviewed text and the registry is its projection" survives any change of version-control system or registry technology. The specific reconciler is disposable; the claim that authorship lives outside the serving system is what keeps corrections honest.

> **Lesson.** If a change to a definition changes the meaning of history, make the definition immutable and versioned, and put the review where the commitment is made rather than where the data is served.

#### ADR-06 · Rolling and calendar windows are both published; the rolling window drives the verdict

**Status:** Accepted  ·  **Shown on views:** 04, 12

*Should the cost of Saturday's outage decay gradually, or reset on the first of the month?*

**Context.** The two window shapes produce opposite incentives over identical arithmetic. A trailing 28-day rolling window never resets: an incident's cost decays minute by minute, which matches how trust in a service actually recovers, and means a team can be frozen for four weeks over one bad Saturday. A calendar month resets on a date: it is comprehensible, it aligns with every other report in the business, and it rewards holding risky changes until the first while making the last week of a bad month unshippable. Both can be computed from the same counters (ADR-01) and both can be published. Only one can drive the freeze, because two verdicts for the same SLO is not a verdict.

**Decision.** Both windows are computed and published for every SLO. The trailing 28-day rolling window is canonical for the verdict and therefore for the freeze; the calendar month is the reporting and compliance window and is the one sealed at close (ADR-14). Every response states which window produced the figure it carries.

**How it is realised on AWS.** budget_window rows carry kind as rolling or calendar with their own open and close timestamps, both reducing over the same sli_bucket rows. The budget projection materialises both. The verdict API reads the rolling figure and names it explicitly in the response; the monthly report and the sealed snapshot use the calendar window. The console shows the two side by side, because a release manager asking "why does it say exhausted" needs to see which window is binding.

| Option | Verdict | Reasoning |
|---|---|---|
| Publish both; rolling drives the verdict | Chosen | Two projections to keep fresh and a reader who must know which is which; no reset to game and a reportable month alongside |
| Rolling only | Rejected | Cleanest incentive; leaves the business with no monthly figure and makes compliance reporting a derived estimate |
| Calendar only | Rejected | Simplest to explain and to seal; creates the end-of-month cliff and the first-of-month release rush |
| Publish both and let each service choose which binds | Rejected | Maximum local autonomy; makes cross-service comparison meaningless and gives every team an argument to have during an incident |

**What it buys**

- No reset date to hold changes against, so the freeze cannot be waited out
- The business keeps a monthly attainment figure that lines up with every other monthly report
- An incident's cost visibly decays, which makes the freeze feel like a constraint rather than a punishment

**What it costs**

- Two figures for the same SLO, and a reader who confuses them will reach the wrong conclusion
- A single bad interval can hold a team for four weeks, which is a real organisational cost and the main source of pressure for an override (ADR-11)
- Twice the materialisation work and twice the recompute on a definition change

**Choose differently when.** If the organisation's reliability commitments were contractual and monthly — credits owed on a calendar month — the calendar window would have to be canonical, because the figure that carries money cannot be the secondary one. In an internal engineering context the rolling window's incentive properties dominate.

**Why it holds up over time.** The arithmetic of both windows is permanent and the choice between them is organisational. What should survive is the structural answer: compute both from the same counters, declare one canonical, and name the window in every response. A platform that publishes one figure without saying which window it is will be misread forever.

> **Lesson.** When two framings of the same measurement produce opposite incentives, publish both and declare which one binds. Hiding one does not remove the argument; it just makes it unanswerable.

#### ADR-15 · Admission refuses an unmeasurable objective, an implicit denominator and unbounded cardinality

**Status:** Accepted  ·  **Shown on views:** 02, 06, 18

*Should the platform accept a definition it knows it cannot measure meaningfully?*

**Context.** Three definitions arrive routinely and all three are worse than useless once published. An objective of 99.999% over a calendar month allows 26 seconds of failure, which minute-grain buckets cannot resolve — the resulting figure is noise presented as precision. A definition that does not say which status classes, synthetic traffic, health checks and client-cancelled requests are excluded has not actually defined anything, and the ambiguity will be resolved later by whoever is under pressure. A definition with unbounded label dimensions is a cardinality incident in the aggregate store, which is the dominant cost driver and the usual cause of a time-series platform becoming unaffordable. In all three cases the cheapest moment to object is before the definition exists.

**Decision.** Admission runs in the reconciler, before anything enters the registry, and refuses three classes of definition: an objective whose budget is finer than the measurement resolution can resolve, a definition that leaves its valid-event denominator implicit, and label dimensions exceeding a declared cardinality limit. A refusal fails the merge with an explanation; it is not a warning.

**How it is realised on AWS.** The reconciler computes the implied budget in events and minutes against the window length and bucket resolution, and refuses when the budget falls below the resolvable floor. The schema requires explicit inclusion or exclusion for each status class and traffic type, so an omission is a parse failure rather than a default. label_dims is a bounded list with a per-SLO cell-count cap that also bounds the outcome cube from ADR-03. All three checks run as part of the pull-request validation, so the feedback arrives in review rather than after merge.

| Option | Verdict | Reasoning |
|---|---|---|
| Refuse at authoring time, as a merge failure | Chosen | Friction at the one moment it is cheap; teams will occasionally be blocked from expressing something they wanted |
| Accept and warn | Rejected | No friction; the warning is read once and the meaningless figure is published for a year |
| Accept, and annotate the resulting figures as low-confidence | Rejected | Honest in principle; nobody reads the annotation and the figure circulates without it |
| Refuse only the cardinality case | Rejected | Protects the cost model, which is the self-interested subset; leaves the two that make the figures wrong |

**What it buys**

- No SLO in the estate carries an objective the data cannot support
- The denominator — where an SLO is really defined and later gamed — is explicit for every SLO by construction
- Storage cost stays bounded per SLO and attributable to its owner

**What it costs**

- A team wanting a very high objective is blocked and may feel the platform is obstructing a legitimate goal
- The resolution check is arithmetic on assumptions; a change in bucket granularity changes which definitions are admissible
- Review-time refusal pushes the conversation into a pull request, where the reasoning is less visible than in a design discussion

**Choose differently when.** If bucket granularity moved from minutes to seconds, the resolvable floor would drop by nearly two orders of magnitude and the first refusal class would mostly disappear — at sixty times the bucket volume. The denominator and cardinality refusals are independent of resolution and would stand regardless.

**Why it holds up over time.** Refusing to represent what you cannot measure is a permanent property of honest measurement systems. The specific thresholds follow the storage design; the principle that admission is a gate rather than a warning is what stops the registry filling with definitions nobody can defend.

> **Lesson.** Validate a definition against the limits of your own measurement before you accept it. A system that will happily store an unanswerable question will be asked it, and will answer.

### Computation and publication

*How the budget is kept fresh at read scale, how a page is decided, and the shape of the document the platform hands over.*

#### ADR-07 · Budget state is materialised incrementally behind a watermark, not recomputed per query

**Status:** Accepted  ·  **Shown on views:** 08, 13, 15

*Can the read path afford the arithmetic the verdict needs?*

**Context.** Computing the budget on every query is the correct-by-construction option: it is always consistent with the current definition, a definition change takes effect instantly, and there is no invalidation logic to get wrong. It is also a scan over up to 40,000 minute buckets per query, and the verdict API is specified at a p99 of 150 ms while absorbing 2,000 requests a second for two minutes during a coordinated release window. Those two facts do not reconcile. Incremental materialisation gives constant-time reads and introduces a staleness window plus invalidation logic that has to stay correct across late data, approved exclusions and retroactive definition changes — three things that all move history.

**Decision.** Budget state is a materialised projection advanced by a watermark over closed minute buckets. Reads are constant time. The projection carries its watermark, every response carries the resulting computation timestamp, and the verdict API returns insufficient-data once staleness exceeds a declared ceiling rather than serving a figure it cannot date. Invalidation is explicit: late data within the horizon, an approved exclusion and a new definition version each enqueue a scoped recompute (ADR-15, view 15).

**How it is realised on AWS.** The budget calculator consumes closed buckets from Data Explorer and writes summed counters plus derived percentages into Azure Cache for Redis, keyed by SLO and window kind, with the watermark minute alongside. The verdict API reads Redis only — it never touches the aggregate store on the request path, which is what makes the read tier independently available (ADR-16). Recompute runs in a separate deployment on preemptible capacity so a full-estate backfill cannot consume the capacity serving verdicts.

| Option | Verdict | Reasoning |
|---|---|---|
| Incremental materialisation with a watermark and a staleness ceiling | Chosen | Constant-time reads and an honest staleness figure; invalidation logic that must survive late data, exclusions and version changes |
| Recompute on read | Rejected | Always consistent and trivially correct; cannot meet 150 ms p99 at 2,000 rps over a 40,000-bucket scan |
| Recompute on read with a short-lived response cache | Rejected | Looks like a compromise; produces the same staleness problem with none of the watermark's visibility into how stale |
| Materialise only the rolling window, recompute calendar on read | Rejected | Halves the projection work; makes the monthly report path a tail-latency risk during exactly the reporting peak |

**What it buys**

- The read path is cheap, flat and independent of history length, so the gate's latency budget is unaffected by retention
- Staleness is a published number rather than an unknown, and the ceiling turns it into a typed refusal
- Recompute is a separate concern on separate capacity, which is what makes backfill affordable

**What it costs**

- Invalidation is the most intricate logic in the platform, and a missed invalidation serves a stale figure that looks authoritative
- A definition change is no longer instantaneous: it is a recompute with a visible duration
- The projection is one more thing that can be wrong in a way the aggregates are not, which is why shadow verification exists (view 15)

**Choose differently when.** If the verdict read rate were low — a handful of queries a day from a human rather than one per deployment from a machine — recompute-on-read would dominate on correctness and simplicity, and the whole projection could be deleted. The decision is driven entirely by the gate being a machine on the request path.

**Why it holds up over time.** The specific store will change. What survives is the pairing: if you cache a derived figure, publish its watermark and refuse to serve past a declared staleness ceiling. A cache without a visible age is the mechanism by which correct systems start giving confidently wrong answers.

> **Lesson.** A derived value served from a cache must carry its own age, and the system must be willing to refuse rather than serve something too old. Freshness you cannot measure is freshness you do not have.

#### ADR-09 · Two multi-window burn-rate rules per SLO, generated from a template rather than authored

**Status:** Accepted  ·  **Shown on views:** 05, 14, 17

*What alerting signal has a defensible relationship to the objective, and how many rules can 1,200 SLOs carry?*

**Context.** Threshold alerting on error percentage is what most organisations have, and it bears no relationship to any commitment: it fires on a bad minute regardless of whether the budget is threatened, and misses a slow leak that quietly exhausts a month. Burn rate fixes the relationship — consumption per unit time, normalised so 1× exhausts the budget exactly at window close — but a single burn-rate rule has to choose between catching a fast outage and catching a slow leak, and cannot do both. The canonical answer is several window and multiplier pairs evaluated simultaneously, which multiplies the rule count across 1,200 SLOs and multiplies the false-positive surface with it. Precision, not recall, is what determines whether pages are still answered in six months.

**Decision.** Each SLO carries exactly two rules, generated from a template rather than authored per service: a fast pair intended to page, and a slow pair intended to raise a ticket. Both require a short confirmation window to agree with the long window before firing. Page-class and ticket-class alerts route to different destinations, and a slow burn never pages. Every evaluation is recorded with its inputs.

**How it is realised on AWS.** Rules are compiled from the registry into an alert_rule row per SLO per class — fast at 14.4× over a one-hour long window with a five-minute confirmation window, slow at 3× over six hours — and evaluated by the burn-rate evaluator on separate ticks. Firing decisions go to Azure Monitor action groups, grouped by common dependency with a declared cap on simultaneous pages. alert_decision rows retain the evaluated inputs for 13 months, which is what makes page precision computable against incident records rather than estimated.

| Option | Verdict | Reasoning |
|---|---|---|
| Two generated pairs per SLO, page and ticket, with confirmation | Chosen | Catches both failure shapes with a bounded rule count; less tuning freedom than per-service authoring |
| A single fast burn-rate rule | Rejected | Half the rules and half the noise; misses the slow leak that exhausts a budget over a week, which is the failure nobody notices |
| The canonical four-window pairing | Rejected | Best recall; 4,800 rules across the estate and a false-positive rate that degrades the precision the pages depend on |
| Per-service authored rules | Rejected | Maximum local fit; 1,200 hand-written configurations drift immediately and cannot be changed estate-wide |

**What it buys**

- A page means the budget is genuinely going, not that a minute looked bad
- Alerting behaviour across 400 services can be changed in one merge, which is the only way the configuration stays coherent
- Page precision is a measured property of the platform rather than an opinion held by the on-call rotation

**What it costs**

- Two rules miss failure shapes a four-window pairing would catch, and that gap is accepted deliberately
- Generated rules fit no service perfectly, so there will be SLOs whose thresholds are visibly wrong for their traffic shape
- Grouping by common dependency can swallow a genuinely separate incident that happens to share one

**Choose differently when.** If the on-call rotation were large enough that page volume were not the binding constraint, the four-window pairing's recall would be worth its false positives. At twelve people carrying 1,200 SLOs it is not — and the moment precision drops far enough that pages are routinely ignored, the platform has failed regardless of how many real failures it detected.

**Why it holds up over time.** Burn rate as the signal is arithmetic and permanent. The rule count is an organisational constraint that will move with the size of the rotation, and the template is the mechanism that lets it move without 1,200 edits. Recording the inputs of every decision is what will still be useful in ten years, when nobody remembers why a threshold was chosen.

> **Lesson.** Alert on the rate at which a commitment is being consumed, not on the symptom. And measure the precision of your own alerts, because an alerting configuration nobody can evaluate will drift until it is ignored.

#### ADR-12 · The platform publishes a signed, typed, time-limited verdict and takes no action

**Status:** Accepted  ·  **Shown on views:** 01, 09, 13, 19

*Should the reliability authority be able to stop a deployment?*

**Context.** If the platform holds the budget and the policy, the shortest path to a real freeze is for the platform to enforce it: call the deployment API, revoke the pipeline's credential, fail the check itself. It removes an integration and makes the policy unambiguous. It also puts this platform on the critical path of every deployment in the organisation, including the rollback that would end an outage, and it makes a false positive in SLI data an estate-wide inability to ship. The alternative is to publish a claim and let the deployment system act on it, which keeps the authority off the critical path but means the policy is only as real as the gate's willingness to honour it.

**Decision.** The platform's output is a signed, typed, time-limited verdict document — healthy, warning, exhausted or insufficient-data — carrying the figure, the definition version, the computation timestamp and the data coverage. The platform blocks nothing, revokes nothing and calls no deployment API. Enforcement is the deployment system's act, performed by reading the verdict and applying its own declared policy.

**How it is realised on AWS.** The verdict API returns a document signed with a non-exportable key in Azure Key Vault Managed HSM, with valid_until set five minutes ahead. The gate authenticates with workload identity federation, verifies the signature itself, and applies the policy recorded against the SLO. Signature validity is short precisely because a signed document is otherwise replayable forever, and the gate's own verification is what makes a cached verdict safe to use during an outage of this platform (ADR-13).

| Option | Verdict | Reasoning |
|---|---|---|
| Publish a signed verdict; the gate enforces | Chosen | One integration and a policy that depends on the gate honouring it; the authority stays off the critical path of the release path |
| The platform calls the deployment API to block | Rejected | Unambiguous enforcement; makes a false positive here an estate-wide inability to ship, including the rollback |
| The platform acts as the CI check itself | Rejected | Neat integration story; the platform's availability becomes the pipeline's availability for every repository |
| Unsigned verdict over mutual TLS | Rejected | Simpler and adequate in transit; a cached verdict is then unverifiable, which removes the fail-static option entirely |

**What it buys**

- An outage of the reliability authority is not simultaneously an observability outage and a deployment outage
- The fail-open or fail-closed decision sits with the system that knows what it is deploying
- A verdict can be cached, archived, attached to a release record and verified later by anyone

**What it costs**

- The policy is only as strong as the gate's implementation, and a bypassable gate makes the arithmetic advisory in fact
- Signing is on the synchronous read path, so HSM latency and availability are inside the verdict API's own budget
- Two systems must agree on the verdict schema, and a schema change is a coordinated release

**Choose differently when.** If deployments were centralised in one pipeline owned by the same team as this platform, direct enforcement would be a smaller risk and would remove a real integration cost. Across 400 services and multiple pipelines it is not, and the rollback argument holds regardless of ownership.

**Why it holds up over time.** "Publish, do not enforce" is the critical design decision restated at the API boundary, and it will be challenged every few years on the grounds that the integration is tedious and the policy feels weak. The counter-argument does not change with scale: an authority on the critical path of the thing it governs cannot fail independently of it.

> **Lesson.** Make the output of a judgement a verifiable document rather than an action. Documents can be cached, audited and acted on by systems you do not control; actions cannot, and they put you on everyone's critical path.

#### ADR-13 · The gate fails static on its cached verdict, and reliability fixes are never frozen

**Status:** Accepted  ·  **Shown on views:** 04, 13, 16, 21

*What does the release gate do when the platform cannot be reached?*

**Context.** Having decided that the gate enforces (ADR-12), the unreachable case needs an answer, and all three available answers are bad. Fail open evaporates the policy during exactly the incident most likely to take both systems down — a regional event, a shared dependency, the 90-minute recompute window after a failover. Fail closed makes this platform a single point of failure for every deployment in the organisation, including the rollback that would end the outage, which inverts the platform's purpose precisely when it matters. Caching the last verdict and continuing to enforce it turns the problem into a staleness ceiling, which is a judgement with no obviously correct value but at least degrades along a measurable axis.

**Decision.** The gate caches the last signed verdict it received and, when the platform is unreachable, continues to enforce that verdict up to a declared staleness ceiling, after which it fails open and records that it did so. Separately and unconditionally, changes classed as reliability fixes or rollbacks are exempt from a freeze by default — a freeze that blocks the fix for the outage that caused it is a defect, not a policy.

**How it is realised on AWS.** The policy object carries gate_mode and an unreachable behaviour, and the verdict's signature plus valid_until let the gate verify a cached document it has held for hours. Rollback and hotfix classes are declared in the policy's exempt_classes and evaluated by the gate from the change's own metadata, not from anything this platform knows. Bypasses and ceiling expiries are reported as first-class figures next to the override register, so a rising bypass rate reads as a policy failure rather than a compliance one.

| Option | Verdict | Reasoning |
|---|---|---|
| Fail static on a verified cached verdict, with a ceiling and fixes exempt | Chosen | Degrades along a measurable axis; the ceiling is a judgement nobody can derive, and the policy weakens with age |
| Fail open immediately | Rejected | Never blocks a fix; removes the policy during the correlated outage that is the likeliest reason the platform is unreachable |
| Fail closed | Rejected | Strongest policy; makes the reliability platform able to stop the rollback that would end an incident |
| Fail closed with a documented manual override | Rejected | Looks balanced; puts a manual step into the incident path at the moment humans are least able to execute one |

**What it buys**

- An outage of this platform never prevents a rollback or a reliability fix
- The policy survives a short outage intact rather than evaporating on the first timeout
- Degradation is observable: bypasses and ceiling expiries are counted and reviewed

**What it costs**

- A gate enforcing an eight-hour-old verdict is enforcing history, and no ceiling value is obviously correct
- The gate must implement signature verification and a cache, which is real work in every pipeline that integrates
- Exempting fixes depends on change classification the gate performs, which can be claimed falsely

**Choose differently when.** If this platform's measured availability were materially better than the deployment system's own, fail closed would stop being a meaningful risk and would give a stronger policy for free. It would still need the rollback exemption, which is the part that is not about availability at all.

**Why it holds up over time.** Fail-static is the standard answer for a control plane whose consumers must keep working without it, and it will outlive the specific ceiling. The rollback exemption is the part most often forgotten and most expensive to learn during an incident; it should be treated as a permanent invariant rather than a configuration choice.

> **Lesson.** Decide what a dependent system does without you before you make it depend on you, and never let your own outage block the action that would fix the thing you are measuring.

### Evidence, privilege and recovery

*What makes the published record trustworthy after the fact, who is allowed to change it, and what happens when the region goes away.*

#### ADR-11 · Exclusions annotate rather than delete, and the SLO's author cannot approve one

**Status:** Accepted  ·  **Shown on views:** 15, 19, 20

*Who is allowed to decide that a bad forty minutes did not count?*

**Context.** Some intervals genuinely should not count: a dependency's declared maintenance, a load test run against production, a measurement artefact already diagnosed. Refusing all exclusions sounds principled and is not — the pressure does not disappear, it relocates, and the next move is to widen the valid-event denominator instead, which is the same adjustment made invisibly. So exclusions have to exist, which creates a privilege whose holder can make any breach appear met. Three mechanisms are available: delete or ignore the interval, annotate it with the original retained, or leave the breach intact and grant a compensating credit. The second preserves the measurement; the third preserves it too but complicates every figure with a second adjustment nobody reads.

**Decision.** An exclusion is an overlay, never an edit: the interval is annotated with a reason code, a named approver, an expiry and a permanent audit record, and the buckets are untouched so the unexcluded figure stays retrievable forever. The author of an SLO cannot approve an exclusion against it, an exclusion beyond 60 minutes requires a second approver, and exclusions have no permanent form — they expire by default at 30 days.

**How it is realised on AWS.** exclusion_window rows in Azure SQL carry the interval, reason_code, two approver fields and expires_at; sli_bucket is never modified. The derivation applies the overlay at window reduction, so both figures are a query away. The privileged API refuses an approval whose actor matches the definition's author, writes the audit ledger entry before the registry write, and counts every exclusion into a register reported weekly alongside attainment. Audit rows are Azure SQL ledger tables, so an after-the-fact edit is detectable rather than merely prohibited.

| Option | Verdict | Reasoning |
|---|---|---|
| Annotate, with separation of duties and an expiry | Chosen | Both figures kept and the adjustment visible; a privilege that still exists and can still be abused by two colluding people |
| Never exclude anything | Rejected | Maximally honest in appearance; relocates the pressure into denominator definitions, where the same adjustment is invisible |
| Delete or ignore the excluded buckets | Rejected | Simplest to implement and to read; destroys the ability to answer what the month looked like before anyone intervened |
| Budget credits instead of exclusions | Rejected | Preserves the measurement perfectly; every figure then carries a second adjustment, and the compound number is one nobody trusts |

**What it buys**

- The question "what did this month look like before the adjustments" always has an answer
- The privilege to improve a figure never sits with the party being measured
- A routine exception becomes visible as a rising count rather than as an unremarkable habit

**What it costs**

- Two colluding approvers defeat the control entirely, and nothing in the architecture detects collusion
- Every published figure now has an excluded and an unexcluded form, and a reader who quotes the wrong one is misleading without intending to
- Exclusion approval is operational work for the reliability lead at exactly the moments they are busiest

**Choose differently when.** If the published figure ever carried contractual weight — customer credits, a regulatory filing — internal separation of duties would stop being sufficient, because the measured party and the approver are both inside the same organisation. At that point the figure needs an external witness: a notarised snapshot or third-party attestation.

**Why it holds up over time.** Separation of duties on an adjustment privilege is structural and survives reorganisations in a way that norms about who may adjust a number do not. The specific thresholds will move; the rule that an author cannot approve their own exclusion is the part worth defending.

> **Lesson.** If a figure can be adjusted, keep the unadjusted figure forever, put the adjustment behind someone other than the party it benefits, and count the adjustments where the figure is published.

#### ADR-14 · At window close an immutable snapshot is sealed and can never be altered

**Status:** Accepted  ·  **Shown on views:** 11, 12, 19

*If last quarter's attainment is quoted in a board pack, can it change afterwards?*

**Context.** Everything else in this platform is deliberately recomputable: a corrected definition yields a corrected history (ADR-05), an approved exclusion changes the current figure (ADR-11), late data moves a bucket within the horizon (ADR-08). That is correct for an operational figure and wrong for a compliance record. A monthly report regenerated from current state will quietly produce a different answer every time it is run, which makes it useless as evidence — and the circumstances under which it changes are exactly the circumstances under which someone would want it to.

**Decision.** At the close of each calendar window, a snapshot of the figure — total, consumed, remaining, coverage and the definition version in force — is sealed in write-once storage with no update or delete path. Subsequent definition changes, exclusions and recomputes produce a new current figure and cannot alter a sealed snapshot. The monthly report is rendered from snapshots, never from current state.

**How it is realised on AWS.** budget_snapshot rows are written once and are immutable in the schema; the rendered snapshot document is written to Azure Blob Storage under a time-based immutability policy with a five-year retention, alongside audit ledger entries in Azure SQL ledger tables. The window manager performs the seal; no other component has write access to that container. A sealed figure that is later found to be wrong is corrected by publishing a dated correction next to it rather than by editing it.

| Option | Verdict | Reasoning |
|---|---|---|
| Seal an immutable snapshot at close; correct by addition | Chosen | A record that means something; a permanent discontinuity in the trend when a long-standing definition error is found |
| Regenerate reports from current state | Rejected | Always reflects the best current understanding; a figure that changes between readings cannot be evidence |
| Snapshot, but allow a privileged correction | Rejected | Handles the genuinely-wrong case; the correction privilege is exactly the one an interested party would want |
| Seal with a short grace period before immutability | Rejected | Pragmatic compromise; the grace period becomes the window in which inconvenient months are revised |

**What it buys**

- A quoted historical figure is stable, dated and attributable to the definition in force at the time
- An auditor's question is answerable without trusting the platform's own access control
- The operational figure can stay freely recomputable, because it is no longer doing the compliance job

**What it costs**

- A definition error found after a close leaves a permanent discontinuity between the sealed figure and the corrected trend
- Write-once retention is irreversible by design: a retention policy set wrongly cannot be corrected, with a five-year blast radius
- Two figures exist for every past month — the sealed one and the current recomputation — and the difference needs explaining

**Choose differently when.** If the figure were purely internal and never quoted outside engineering, regenerating from current state would be simpler and the loss would be theoretical. The moment a number appears in a board pack or a customer conversation, it needs to be the same number next quarter.

**Why it holds up over time.** Append-only evidence with correction-by-addition is how accounting has worked for centuries and how every durable audit system works. It will outlive the storage product; the thing to preserve is that the component which seals is the only one with write access, and that nothing has an update path.

> **Lesson.** Separate the operational figure from the record. The first should be freely recomputable; the second should be sealed, dated, and corrected only by publishing a correction beside it.

#### ADR-16 · Derived state is rebuilt rather than restored, and there is one write region

**Status:** Accepted  ·  **Shown on views:** 11, 15, 16

*What is actually worth backing up, and can two regions compute the same budget?*

**Context.** The platform holds four kinds of state with very different properties. The registry is irreplaceable and small. The minute buckets are irreplaceable, large and replayable from quarantine. The sealed snapshots and audit ledger are evidence with no update path. The budget projection, compiled rules and report extracts are none of these: they are a deterministic function of the first two, and a backup of them is a slower way to obtain a staler version of something that can be recomputed. Separately, there is a question about regions: a second write region would improve availability and would also have two independent computation planes deriving the same budget from partially overlapping, independently late bucket sets — producing two figures that no key reconciles, in a system whose entire purpose is to produce one.

**Decision.** Derived state is not backed up at all; its recovery objective is stated as recompute time — a trustworthy projection within 90 minutes of regional recovery. The registry and audit ledger are backed up with a one-minute RPO and geo-replicated; aggregates are replicated and replayable; snapshots are write-once with cross-region redundancy. There is exactly one write region. Failover is promote-then-recompute, and the honest intermediate state is insufficient-data (ADR-04).

**How it is realised on AWS.** Azure SQL is zone-redundant with point-in-time restore and an active geo-replica in the secondary; the Data Explorer cluster is followed in the secondary; snapshots and audit sit in read-access geo-redundant immutable storage. Redis holds the projection and is not backed up — the recompute runner rebuilds it from aggregates plus registry. The secondary region's compute is scaled to zero. A shadow rebuilder recomputes sampled SLOs continuously and alerts on divergence, because a recovery path exercised only during an incident is a claim rather than a capability.

| Option | Verdict | Reasoning |
|---|---|---|
| One write region; rebuild derived state rather than restore it | Chosen | One answer always, and a recovery story that improves as compute gets cheaper; 90 minutes of insufficient-data after a regional loss |
| Two active write regions | Rejected | Best availability; two computation planes produce two irreconcilable budgets for the same SLO |
| One write region, with the projection backed up for fast restore | Rejected | Shortens the gap; restores a figure that is stale by exactly the backup interval, which the verdict would have to disclose anyway |
| Active-passive with synchronous projection replication | Rejected | Near-zero gap; couples the read path's latency to cross-region replication for state that is disposable by construction |

**What it buys**

- There is never more than one budget figure for an SLO
- Backup scope is small, so restore drills are cheap and actually get run
- The recovery story improves automatically as compute gets faster, where a restore of growing data gets slower

**What it costs**

- A 90-minute window of insufficient-data after a regional loss, during which every gate falls back to its unreachable behaviour (ADR-13)
- The secondary's capacity on promotion is a cold-start assumption until a game day proves it
- Recompute capacity has to be available during a regional incident, which is when it is most contended

**Choose differently when.** If the verdict API ever became a hard dependency of something that cannot tolerate 90 minutes of insufficient-data — an automated capacity controller, a customer-facing commitment — a second write region would become necessary, and the reconciliation problem would have to be solved rather than avoided, probably by partitioning SLO ownership by region.

**Why it holds up over time.** Classifying state by rebuildability rather than by importance is the decision that keeps paying off: it shrinks the backup surface, makes recovery testable, and improves with hardware. The one-write-region constraint is specific to needing a single answer and would only change if the arithmetic could be partitioned.

> **Lesson.** Before backing something up, ask whether it can be recomputed. If it can, spend the effort on making the recomputation fast and continuously exercised instead of on storing a stale copy.

#### ADR-17 · An independent watchdog in a third region holds the platform's own availability signal

**Status:** Accepted  ·  **Shown on views:** 16, 17, 21

*Who notices when the thing that notices everything stops working?*

**Context.** The platform computes SLOs, so the obvious way to monitor it is to give it SLOs of its own and compute them with the same engine. That is useful and insufficient: the measurement is circular, and the failures that matter most are exactly the ones that take the engine down with the signal. A platform that is up, fast and serving stale or wrong figures passes every internal check, and a platform that is entirely down produces no alert at all because the thing that would produce it is the thing that is down. Nothing inside the system can close that gap, and a self-check deployed inside the same region and the same subscription shares enough failure modes to be nearly as blind.

**Decision.** The platform holds SLOs on itself — budget freshness, verdict availability, alert detection latency — computed by the same engine for day-to-day operations. Independently, a deliberately minimal watchdog deployed in a third region probes the verdict API from outside and alerts on absence and on staleness, with no dependency on any component it is watching. The package documents that the platform is not the sole authority on its own availability.

**How it is realised on AWS.** An Azure Functions app in a third region, on its own subscription and its own alerting path, calls the verdict API on a fixed interval with a synthetic read-only SLO, verifies the signature, and checks the computation timestamp against its own clock. It shares no store, no identity and no deployment pipeline with the platform. It raises to the platform rotation through an action group that does not traverse the primary region. Its own availability is explicitly not measured by this platform, which bounds the recursion at one level.

| Option | Verdict | Reasoning |
|---|---|---|
| Self-measured SLOs plus a minimal independent watchdog elsewhere | Chosen | Bounds the circularity for the cost of a second deployment; the recursion is stopped by fiat rather than solved |
| Self-measured SLOs only | Rejected | No extra component; produces no signal in the one failure mode where a signal is essential |
| A full second instance of the platform watching the first | Rejected | Rich signal and genuine independence; doubles the operational surface and raises the same question one level up |
| Rely on the consumers to notice | Rejected | Zero cost; the gate failing static is silent by design, so nobody notices until a release is questioned |

**What it buys**

- Total unavailability of the platform produces an alert from outside it
- Staleness is checked against an independent clock, so a stuck computation plane serving old figures is caught
- The watchdog is small enough to be obviously correct, which is the only useful property for a component of this kind

**What it costs**

- A second deployment, subscription and alerting path to maintain, which will be neglected because nothing ever happens to it
- The watchdog's own availability is unmeasured, so the recursion is stopped by assertion
- A correlated failure of platform and watchdog is undetected until a human notices

**Choose differently when.** If an external synthetic monitoring service were already trusted for estate-wide availability checks, it would be the better watchdog — genuinely outside the organisation's cloud estate, and already staffed. Building one is only justified because the probe needs to verify a signature and inspect a computation timestamp rather than just receive a 200.

**Why it holds up over time.** The need for an external observer of an observer is permanent and the answer is always the same shape: something small, something elsewhere, something with no shared dependencies, and an explicit decision about where the recursion stops. The implementation will be replaced; the reasoning will not.

> **Lesson.** Any system that reports on others needs something outside it reporting on it, and that something should be small enough that its own correctness is obvious rather than measured.

## Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on service-level objectives, and the difference matters when reading the decision records.

| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Outcome cube | The small declared set of counter cells a source emits per SLO per minute — counts by status class and by latency bucket. | The emission contract, and the thing that decides which future predicate changes are recomputable (ADR-03). | "Metrics", which implies the platform receives arbitrary time series rather than a bounded, declared shape. |
| Valid-event denominator | The explicit statement of which events count toward the SLO at all — which status classes, synthetic traffic, health checks and client-cancelled requests are in or out. | Where an SLO is really defined, where it is later quietly widened, and the one thing admission refuses to let a definition leave implicit (ADR-15). | "Total requests", which sounds unambiguous and is the source of most disagreement about what an SLO means. |
| Coverage | The fraction of a window's expected minute buckets that were actually present and not no-data. | Published alongside every figure, so a number is never read as more certain than the data behind it; below the floor the verdict becomes insufficient-data (ADR-04). | "Data quality", which bundles completeness together with correctness and hides which one failed. |
| no-data | A bucket state meaning no valid events were observed for that minute — neither good nor bad, and excluded from both numerator and denominator. | The mechanism that stops a stopped metrics pipeline reading as a perfect month. | Zero, which an arithmetic pipeline will divide by or treat as success. |
| insufficient-data | A first-class verdict value, returned instead of a figure when coverage is below the floor or the projection is staler than the ceiling. | The platform's way of declining to publish something it cannot substantiate; every consumer must handle it as distinct from healthy. | An error or a null, both of which consumers routinely treat as "carry on". |
| Burn rate | Error budget consumption per unit time, normalised so that 1× exhausts the window's budget exactly at its close. | The only alerting signal with a defensible relationship to the objective (ADR-09). | Error rate or error percentage, which fires on a bad minute regardless of whether any commitment is threatened. |
| Confirmation window | A short evaluation window that must agree with the long window before an alert fires. | What stops a single anomalous minute paging an on-call engineer. | "For duration" or a debounce, which delays an alert without requiring a second independent agreement. |
| Verdict | A signed, typed, time-limited document stating healthy, warning, exhausted or insufficient-data, with the figure, definition version, computation timestamp and coverage. | The platform's entire output to machines. It is a claim, not an instruction (ADR-12). | "Gate result" or "check status", which imply the platform performed the enforcement. |
| Watermark | The most recent minute the budget projection has incorporated. | What makes staleness measurable, and therefore what makes the staleness ceiling and the insufficient-data refusal possible (ADR-07). | "Last updated", which is usually the time of the write rather than the time of the data. |
| Exclusion window | An approved interval annotated as outside the valid-event denominator, with a reason code, two approvers and an expiry. | The sanctioned way to say a bad interval should not count, designed so the unexcluded figure survives (ADR-11). | "Maintenance window" or "downtime exemption", which imply the data was removed rather than annotated. |
| Sealed snapshot | The immutable record of a calendar window's figure, written once at close with no update path. | The compliance record, as distinct from the freely recomputable operational figure (ADR-14). | "Monthly report", which in most systems is regenerated from current state and therefore changes between readings. |
| Fail static | The release gate continuing to enforce its last verified cached verdict while the platform is unreachable, up to a declared staleness ceiling. | The declared answer to the unreachable case, chosen over fail-open and fail-closed (ADR-13). | "Fail safe", which is ambiguous here — open and closed are each the safe option under a different failure. |
| Derived state | The budget projection, compiled alert rules and report extracts — a deterministic function of the registry and the minute buckets. | Deliberately not backed up; its recovery objective is recompute time rather than an RPO (ADR-16). | "Cache", which understates it — consumers read it as authoritative, so its staleness has to be published. |
| Measurement failure | An alert class meaning the platform cannot see an SLO, as distinct from an alert meaning the service is failing. | Routed to the platform rotation rather than the service rotation, so triage starts at the broken system (ADR-10). | A low-severity reliability alert, which still points the reader at the wrong system. |
