Incident Management Platform · View 21 of 34 · 5 · Runtime
Decisions
- Load is shed per integration at the edge, in HAProxy stick tables keyed by integration credential. The noisy integration gets 429s with a retry-after; every other integration is admitted at its normal rate.
- Everything admitted is buffered, grouped and counted. Flood control attaches alerts to the open incident as counts rather than one event each, so the timeline shows 2.4 million alerts in one line and ClickHouse holds every one of them.
- Paging capacity is reserved in CPU, not hoped for. The dispatcher's systemd slice has guaranteed cores that the gateway's slice cannot borrow.
Targets
- 10× burst for 120 s with no increase in time to first notification for incidents already open. No alert dropped after admission. Ingest consumer lag bounded and alarmed at 60 s.
Proof
- A recorded storm is replayed against the staging cells on every paging-path release (view 27). The pass condition is not throughput; it is the dispatch latency of an unrelated SEV2 injected in the middle of the storm.