Webhook Delivery Service · View 15 of 20 · Operations
Decisions
- Delivery workers live in egress-only subnets with no route into the capture, control or product networks. That topology is the control that holds when the address guard has a bug (ADR-12).
- Both regions' NAT ranges are published from day one. Otherwise a failover is also an allowlist change for 40,000 consumers, which is a failover nobody will ever perform (ADR-11).
- Standby is warm, not active-active: intake at a minimum task floor, delivery workers scaled to zero, state replicated continuously. Delivery RTO ≤ 15 minutes is cheaper to buy this way than active-active is to operate.
- DynamoDB global tables and S3 cross-region replication carry the RPO. Accepted events are RPO 0 in-region and ≤ 5 s cross-region.
Assumptions
- Three availability zones per region; SQS, DynamoDB, S3 and KMS are taken as regionally available services with their own multi-AZ properties.
- Delivery RTO ≤ 15 min, intake RTO ≤ 5 min, subscription store RPO ≤ 5 s / RTO ≤ 10 min.
- Duplicates after a regional failover are expected and are covered by the at-least-once contract rather than prevented.
Risks
- NAT gateway port exhaustion presents as connection failures against perfectly healthy consumers, which is the hardest failure on this platform to diagnose. Egress capacity is tracked as a scaling dimension with its own headroom alarm.
- Multi-region delivery is Phase 2. Single-region is the MVP, and a regional loss in the MVP is a delivery outage bounded by the retry window rather than a failover.