Incident Management Platform · View 27 of 34 · 6 · Operations
Decisions
- The outpost cell is changed first, because it is the cell whose failure is best covered by the other two during normal operation. Cell B follows the next day and cell A the day after (ADR-31).
- No paging-path change starts while a SEV1 or SEV2 is open, and none starts outside hours when both the platform team's regions are awake. A deploy is the most common cause of an outage, and this platform's outage is everyone's.
- Binaries, configuration and tzdata are staged on every host before the switch. A deploy needs no registry, Git or DNS reachable at the moment it happens, and rollback is a symlink and a restart.
Gates
- Schedule resolution for every rotation across every DST transition to 2030. Replay of a recorded storm with an injected SEV2 dispatched on time. Three synthetic pages acknowledged through the new cell before the next cell is touched.
Control plane
- Control-plane services deploy continuously through Argo CD with ordinary canaries. Their failure costs the console for minutes, which the degradation order allows.