Incident Management Platform · View 09 of 34 · 3 · Structure
Decisions
- Paging services run as static Go binaries under systemd on hosts the platform team owns. The estate's Kubernetes clusters are among the things this platform pages for, so the paging path cannot live on them (ADR-02).
- Ingest and paging share hosts but not CPU. The gateway runs in a systemd slice with a hard CPU quota; the engine, timers and dispatcher run in a slice with reserved CPU. A storm raises ingest lag; it cannot delay a dispatch.
- Two JetStream domains per cell: a site-local ingest domain replicated across the cell's three hosts, and a paging domain whose replicas sit one per site (ADR-07). Bulk alert traffic never crosses a site link.
Targets
- Any one host of three may be lost with no effect. A whole cell may be lost and the anycast address moves to another cell within 2 minutes, the paging path RTO. Timer shards owned by the lost cell are taken over within 3 seconds.
Risks
- A cell is also a place where a bad deploy lands on every service at once. Deploys go one cell at a time, one day apart, with a synthetic page acknowledged between them (view 27).