Health Check & Service Discovery  ·  View 17 of 21  ·  Operations

Observability

Six signal families across five stages, reduced to four alarms that map to four distinct actions.

Editable source SVG draw.io All views
Collection Evaluation Propagation Client Outcome Freshness Sample age p99 Decision age Propagation delay p99 View staleness Convergence lag Correctness Unattributed signals Unknown-state count Version regressions Cache-serve ratio Black-holed requests Stability Probe flap rate Transitions per instance Delta churn volume Ejection rate Retry rate by endpoint Capacity Results per second Shard evaluation lag Open subscriptions Client CPU share Eligible fraction Safety Prober reachability Fraction-guard trips Shed episodes Shrink-cap hits Static-serving minutes Governance Contract changes staged Manual overrides Library version spread Removal-drill results Observability — Six Signal Families Across Five Stages Four alarms, not thirty: withdrawal delay past budget, fraction-guard trip, version regression, and static-serving minutes above zero when nothing is declared broken. v 1.0 · owner Reliability Architecture

The four alarms

  • Withdrawal propagation delay past its budget — the data plane is more wrong than the design allows.
  • Fraction guard tripped — either a fleet is broken or a health check is, and the two need opposite responses.
  • View version regression at a client — a correctness bug, not a capacity one.
  • Static-serving minutes above zero with nothing declared broken — clients are routing on stale truth and nobody asked them to.

Assumptions

  • Convergence lag and eligible fraction are reported per service, not per platform; a platform average hides every real event.

Gap

  • Client library version spread is the signal most likely to be missing in practice, and the one that decides whether a client-side guard actually exists everywhere.