Incident Management Platform · View 02 of 34 · 1 · Context and scope
Decisions
- The paging path and the incident record are separate planes with separate availability targets: 99.99% for the path that wakes people, 99.9% for the console, schedules and reviews. Nothing in the first five stages calls anything in the sixth.
- Schedules are resolved ahead of time by the control plane and pushed into the paging path as a snapshot. The paging path reads who is on call; it never computes it.
- The notification ledger is written inside the paging path, next to the dispatcher, so the evidence of what was sent survives an outage of the incident store.
Targets
- Alert receipt to first notification dispatched: p95 ≤ 15 s, p99 ≤ 30 s. Acknowledgement to escalation cancelled: p99 ≤ 5 s. Coverage snapshot staleness after a schedule change: p99 ≤ 60 s.
Cost accepted
- Two representations of on-call state and a window where they disagree. The failure this allows is paging the person who was on call a minute ago. The failure it prevents is paging nobody because the schedule database was down.