Event-Driven Notification Platform · View 20 of 26 · 5 · Operations
Autoscaling, Capacity and Backpressure
What each tier scales on, where scaling stops, and what the platform does once it has stopped.
Copy
PNG
PDF
⋯
Editable source
SVG
draw.io
All views
Pressure signal
Pressure signal
Consumer lag
per topic per tier
Consumer lag...
Due-work depth
scheduled backlog
Due-work depth...
Provider latency
p95 per vendor
Provider latency...
Controller
Controller
KEDA scaler
lag-driven
KEDA scaler...
HPA
CPU for stateless APIs
HPA...
Flink autoscaler
reactive parallelism
Flink autoscaler...
Scaled unit
Scaled unit
Ingest API
12 to 90 pods
Ingest API...
Render Service
20 to 160 pods
Render Service...
Channel workers
80 to 600 pods
Channel workers...
Flink task managers
24 to 120
Flink task managers...
Hard ceiling
Hard ceiling
Partition count
concurrency ceiling
Partition count...
Provider TPS
contracted, not elastic
Provider TPS...
Database connections
PgBouncer pool
Database connections...
Backpressure
Backpressure
Shed P2 first
marketing pauses
Shed P2 first...
429 with Retry-After
at the gateway
429 with Retry-After...
Park to DLQ
last resort, replayable
Park to DLQ...
ceiling reached
ceiling reached
vendor cap
vendor cap
still over budget
still over budget
Autoscaling, Capacity and Backpressure
Autoscaling, Capacity and Backpressure
Application we own
Application we own
Security / platform
Security / platform
Queue / topic
Queue / topic
External / third party
External / third party
Data store
Data store
Decision point
Decision point
Interface / broker
Interface / broker
failure / alternate
failure / alternate
Capacity target: 5 000 events/s normal, 25 000 peak, fan-out 1.8, so 45 000 notifications/s peak and roughly 750 M a day. Scaling stops at the partition count, which is why partitions are provisioned for peak and not for today.
Capacity target: 5 000 events/s normal, 25 000 peak, fan-out 1.8, so 45 000 notifications/s peak and roughly 750 M a day. Scaling stops at the partition count, which is why partitions are provisioned for peak and not for today.
v 1.0 · owner Data & AI Global Practice · date 2026-08
v 1.0 · owner Data & AI Global Practice · date 2026-08
Text is not SVG - cannot display
Capacity targets
5 000 events/s normal, 25 000 peak, sustained for 30 minutes
Fan-out 1.8, so 45 000 notifications/s at peak and about 750 M a day
Maximum event size 64 KB hard, 4 KB at p99 — events carrying more than that are carrying data that belongs behind a reference
Where scaling stops
Partition count is the concurrency ceiling, which is why partitions are provisioned for peak rather than for today
Provider TPS is contracted, not elastic — beyond it, queueing is the only option
Database connections are bounded by PgBouncer, deliberately, so a scale-out event cannot exhaust PostgreSQL
Backpressure order
Shed bulk traffic first — marketing pauses before anything transactional degrades
Then 429 with Retry-After at the gateway, per tenant, so the producer learns rather than times out
Park to the DLQ only as a last resort, because everything parked has to be replayed and deduplicated later
◀ Delivery Pipeline and Environments
All views
Observability and Operations ▶