Incident Management Platform · View 28 of 34 · 6 · Operations
Decisions
- A synthetic alert enters every cell every 60 seconds, becomes a test incident, pages a platform-owned SIM and a test app, and is acknowledged by a robot. Its end-to-end time is the platform's most honest SLI.
- The dead-man switch lives at site C and expects a heartbeat from every cell and from the synthetic loop. Silence pages the platform rotation through carrier B directly, bypassing the paging services entirely (ADR-30).
- Freshness is measured separately from volume: snapshot age, timer drift, reconcile lag and projection age. A platform can be fully up and paging from a two-hour-old schedule.
Stack
- Prometheus per cell with Alertmanager, Loki for logs, OpenTelemetry collectors with bounded buffers, Grafana for dashboards. The observability stack is not on the paging path: losing it loses sight, not pages.
Risks
- A platform that pages about itself through itself fails silently in exactly the case that matters. Hence two independent routes for its own alerts, and a quarterly drill that kills the paging services on purpose.