SLO and Error Budget Service  ·  View 08 of 21  ·  Structure

Platform Components

The containers in one region, and the three edges that carry the whole verdict story.

Editable source SVG draw.io All views
Microsoft Azure — primary region Reader surface (Container Apps behind Front Door and API Management) Reliability console Verdict API signed, typed Budget query API Report renderer monthly Control plane (Container Apps) Definition reconciler Admission checks Policy engine Exclusion workflow Computation plane (Container Apps, KEDA-scaled) Event-time bucketer Budget calculator Window manager Burn-rate evaluator Recompute runner spot pool State SLO registry Azure SQL SLI aggregates Data Explorer Budget projection Cache for Redis Window snapshots immutable Blob Audit ledger SQL ledger tables Indicator stream Event Hubs Definitions repository Release gate materialise read verdict? Platform Components — Container View Application we own Interface / broker Decision point Security / platform Data store Queue / topic External / third party synchronous Four tiers in one region, carrying only the three edges this view alone can show: a gate asks, the verdict API reads the projection, the calculator writes it. Nothing else a reader depends on is written by the computation plane. The inbound stream and the definitions repository are drawn in views 09 and 10, the signing path in view 20, and the secondary region in view 16. v 1.0 · owner Reliability Architecture · date 2026-10

Decisions

  • Four tiers deploy independently on one container platform. The computation tier can be scaled, drained or rolled back without touching the read path, which is what keeps a backfill off the gate's latency budget (ADR-07).
  • The recompute runner is a separate deployment on preemptible capacity rather than a mode of the calculator, so an expensive full-estate rebuild cannot consume the capacity serving verdicts (ADR-16).
  • Exactly one component writes the registry and exactly one signs a verdict. Both are named here and traced in view 20 (ADR-12).

Assumptions

  • Managed PaaS throughout — no Kubernetes to operate, on the view that a reliability platform should not be the most operationally demanding thing in the estate.
  • The budget projection is a cache, so losing the whole of it is a 90-minute recompute rather than an incident with data loss.

Deliberately omitted

  • The ingest chain and the definitions path, drawn in full in views 09 and 10.
  • The signing call to the HSM, drawn in view 20 where the privilege story makes it legible.