Incident Management Platform  ·  View 28 of 34  ·  6 · Operations

Observability — Watching the Watcher From Outside

What is measured at each stage of the loop, which signals are about freshness rather than volume, and how a silent platform is detected by something that is not the platform.

Editable source SVG draw.io All views
Ingest Decide Dispatch Acknowledge Record Metrics Accepted · shed Suppression rate Receipt latency Time to ack Reconcile lag Freshness Consumer lag Snapshot age Timer drift p99 Cancel latency p99 Projection age Synthetic Test alert · 60 s Test incident opened SMS and call to own SIM Robot presses 1 Event appears in log Outside view Carrier B only Dead-man switch Observability — Watching the Watcher From Outside The platform's own alerts are paged through itself and, independently, by the outpost watchdog on carrier B. A silent platform is the alert. v 1.0 · owner Reliability Architecture · date 2026-09

Decisions

  • A synthetic alert enters every cell every 60 seconds, becomes a test incident, pages a platform-owned SIM and a test app, and is acknowledged by a robot. Its end-to-end time is the platform's most honest SLI.
  • The dead-man switch lives at site C and expects a heartbeat from every cell and from the synthetic loop. Silence pages the platform rotation through carrier B directly, bypassing the paging services entirely (ADR-30).
  • Freshness is measured separately from volume: snapshot age, timer drift, reconcile lag and projection age. A platform can be fully up and paging from a two-hour-old schedule.

Stack

  • Prometheus per cell with Alertmanager, Loki for logs, OpenTelemetry collectors with bounded buffers, Grafana for dashboards. The observability stack is not on the paging path: losing it loses sight, not pages.

Risks

  • A platform that pages about itself through itself fails silently in exactly the case that matters. Hence two independent routes for its own alerts, and a quarterly drill that kills the paging services on purpose.