Incident Management Platform · View 01 of 34 · 1 · Context and scope
Decisions
- The platform consumes alerts and never evaluates metrics or owns alerting rules. A monitoring system that is wrong about a threshold is fixed where the threshold lives; this platform decides only who hears about it.
- Carriers, Apple and Google push are drawn outside because they are. No on-premises design owns the last mile to a phone. What the design owns is the diversity: two carrier routes with separate contracts and separate network paths, and never push as the only channel for SEV1 or SEV2.
- The corporate identity provider and the delivery pipelines sit in the bottom row on purpose. The platform uses both and depends on neither to page: acknowledgement needs no login, and a missing change feed costs context, not a page.
Assumptions
- One organisation, 400 responders, 40 services, 25 rotations, three timezones, 5,000 incidents and 60,000 notifications a month. All figures are the requirement's stated assumptions.
- The requirement names Google Cloud with an out-of-cloud notification provider. This design runs the same problem on hardware the organisation owns, with open-source software only; the mapping is in the README.
Deliberately out
- Metric evaluation, alert rule authoring, status-page hosting for external customers beyond a static page, and chat itself. Mattermost and the issue tracker are integrations, not components.