SLO and Error Budget Service · View 16 of 21 · Operations
Decisions
- One write region, because two regions computing the same budget from partially overlapping buckets produce two figures that no key reconciles — and a reliability authority cannot have two answers (ADR-16).
- Recovery is promote-then-recompute, targeting 90 minutes to a trustworthy projection rather than seconds to a stale one. The honest intermediate state is insufficient-data (ADR-04).
- The watchdog runs in a third region and is deliberately minimal. A platform cannot be the sole authority on its own availability, and a watchdog that shares the estate's failure modes measures nothing (ADR-17).
Targets
- Control plane ≥ 99.9% monthly, verdict API ≥ 99.95% monthly, alert evaluation ≥ 99.9% monthly — all assumed.
- Registry RPO ≤ 1 minute via geo-replication; derived state restored by recompute within 90 minutes — assumed.
- Zone-redundant across three availability zones for every stateful component in the write region.
Risks
- A 90-minute window of insufficient-data after a regional loss means the release gate falls back to its declared unreachable behaviour during exactly the period the organisation is most likely to be shipping fixes (ADR-13).
- The secondary is scaled to zero, so its capacity on promotion is a cold-start assumption rather than a measured one until a game day proves otherwise.