Health Check & Service Discovery  ·  View 11 of 21  ·  Data

State Classes

Three classes ordered by what happens if the data is lost. Only one of them has an RPO worth arguing about.

Editable source SVG draw.io All views
Desired state — RPO 0, system of record Registry (regional, strongly consistent) Services 450 rows Instances + leases 60,000 rows Health contracts versioned Policy Topology + domains Budgets + fractions Resolution authz Observed state — RPO 60 s, recomputable Samples Probe results 7 days Caller outcomes 7 days Lease renewals rolling Evidence kept longer Transitions + cause 90 days Audit records 5 years, immutable Computed state — no RPO, rebuilt not restored Control plane memory Eligibility + weights Versioned views Client memory and disk Last-known-good cache per client, durable rebuild input rebuild input pushed State Classes — Ordered By What Happens If It Is Lost Data store Application we own synchronous event / async The one durable copy on the request path is the client cache — and it is the one copy the platform cannot restore. v 1.0 · owner Reliability Architecture

Decisions

  • Desired state is the system of record, RPO 0. Observed state is recomputable, RPO 60 s. Computed state has no RPO because it is rebuilt (ADR-02).
  • Published views and client subscriptions are held in memory and are never durable artefacts of the control plane.
  • The one durable copy on the request path is the client cache — and it is the one the platform cannot restore.

Assumptions

  • Raw samples 7 days; transitions with evidence 90 days; registration, policy and evacuation audit 5 years, immutable.
  • RTO 5 minutes to restore a regional control plane including a full rebuild of computed views.

Consequence

  • Rebuild is exercised on a schedule. A rebuild path only tested during an incident is a hope with a runbook.