SLO and Error Budget Service  ·  View 01 of 21  ·  Context and scope

System Context

What is inside the boundary, who touches it, and the two things the platform refuses to own.

Editable source SVG draw.io All views
Measurement plane — not ours Azure Monitor Managed Prometheus Synthetic probes People Service team lead Release manager Platform SRE Reliability lead Who acts on the verdict CI/CD release gate On-call paging Incident platform Executive reporting SLO & Error Budget Service computes the verdict; owns no signal Platform dependencies Definitions repository Microsoft Entra ID Key Vault Managed HSM declares asks the gate on call owns policy verdict pages correlates monthly buckets rules probes SLO as code identity signing SLO and Error Budget Service — System Context Security / platform External / third party Person or role synchronous event / async batch two-way In scope: deriving an error budget from someone else's aggregates and publishing a signed verdict. Out of scope: collecting telemetry, storing traces, drawing dashboards, and blocking a deployment. The audit ledger and key custody are omitted here and appear in views 11, 19 and 20. v 1.0 · owner Reliability Architecture · date 2026-10

Decisions

  • The platform's output is a derived figure and a signed verdict — never a collected metric and never an enforcement action. Both the measurement plane above and the release gate to the right are outside the boundary (ADR-02).
  • The release gate is modelled as an actor rather than an integration, because it is the only consumer that acts on a verdict with no human reading it. That is what forces the verdict to be signed, typed and time-limited (ADR-12).
  • The definitions repository is an inbound dependency: the registry is a projection of reviewed text, not a system of authorship (ADR-05).

Assumptions

  • 1,200 SLOs across 400 services and 60 critical user journeys at launch, growing 30% a year.
  • Upstream aggregation reduces 2.5 million requests/second of measured traffic to ≤ 25,000 samples/second entering the platform.
  • Both the incident platform and the paging system already exist and are owned elsewhere.

Deliberately omitted

  • Key custody and the audit ledger, which would crowd this view; they are drawn in views 19 and 20.
  • Dashboards and trace storage, which belong to the observability platform and are not a gap in this design.