SLO and Error Budget Service  ·  View 16 of 21  ·  Operations

Deployment Architecture

One write region across three zones, and a recovery that is a recompute rather than a failover.

Editable source SVG draw.io All views
Microsoft Azure — primary write region Zone-redundant across three availability zones Container Apps env all tiers, KEDA Azure SQL zone-redundant, PITR Data Explorer 3-node cluster Cache for Redis zone-redundant Regional shared services Front Door global anycast API Management Event Hubs zone-redundant Key Vault Managed HSM Microsoft Azure — secondary region (warm, no writes) Registry geo-replica read-only Data Explorer follower aggregates Container Apps scaled to zero Snapshots + audit RA-GRS, immutable Third region — independent watchdog Platform watchdog Functions, own SLOs geo-replicate probes the verdict API Deployment — One Write Region, Recompute-Based Recovery Application we own Data store Interface / broker Queue / topic Security / platform event / async two-way One write region, because two regions computing the same budget from partially overlapping buckets produce two figures that no key reconciles. The aggregate cluster is followed in the secondary alongside the registry replica. Failover is a promotion plus a recompute, targeting 90 minutes to a trustworthy projection rather than seconds to a stale one. The watchdog is deliberately elsewhere: a platform cannot be the sole authority on its own availability. v 1.0 · owner Reliability Architecture · date 2026-10

Decisions

  • One write region, because two regions computing the same budget from partially overlapping buckets produce two figures that no key reconciles — and a reliability authority cannot have two answers (ADR-16).
  • Recovery is promote-then-recompute, targeting 90 minutes to a trustworthy projection rather than seconds to a stale one. The honest intermediate state is insufficient-data (ADR-04).
  • The watchdog runs in a third region and is deliberately minimal. A platform cannot be the sole authority on its own availability, and a watchdog that shares the estate's failure modes measures nothing (ADR-17).

Targets

  • Control plane ≥ 99.9% monthly, verdict API ≥ 99.95% monthly, alert evaluation ≥ 99.9% monthly — all assumed.
  • Registry RPO ≤ 1 minute via geo-replication; derived state restored by recompute within 90 minutes — assumed.
  • Zone-redundant across three availability zones for every stateful component in the write region.

Risks

  • A 90-minute window of insufficient-data after a regional loss means the release gate falls back to its declared unreachable behaviour during exactly the period the organisation is most likely to be shipping fixes (ADR-13).
  • The secondary is scaled to zero, so its capacity on promotion is a cold-start assumption rather than a measured one until a game day proves otherwise.