Incident Management Platform · View 19 of 34 · 5 · Runtime
Decisions
- The gateway answers 202 with an alert id as soon as the message is durably in the ingest domain. The source's retry logic is then only exercised by genuine failures, and a retry that does arrive is absorbed by deduplication.
- The timer is written before the first send. If the dispatcher crashes between arming and dispatching, the timer fires on another owner and the page still goes out, possibly twice, which the human-side dedup and Principle 3 both accept.
- Acknowledgement is a compare-and-set on incident state in the paging domain. The timers and the dispatcher check that state before every fire and every send, so an acknowledgement stops queued work as well as future work.
Targets
- Receipt to first dispatch: p95 ≤ 15 s, p99 ≤ 30 s. Keypress to escalation cancelled: p99 ≤ 5 s. Timer accuracy: ± 5 s at p99, including across a process restart.
Channel policy shown
- SEV1: push immediately, SMS at 30 s, voice at 2 min, all inside a 5-minute step. SEV2: push and SMS immediately, voice at 5 min, 15-minute step. SEV3 and SEV4: push and email, no voice, which is where most of the spend is saved (view 29).