Event-Driven Notification Platform  ·  View 20 of 26  ·  5 · Operations

Autoscaling, Capacity and Backpressure

What each tier scales on, where scaling stops, and what the platform does once it has stopped.

Editable source SVG draw.io All views
Pressure signal
Pressure signal
Consumer lag
per topic per tier
Consumer lag...
Due-work depth
scheduled backlog
Due-work depth...
Provider latency
p95 per vendor
Provider latency...
Controller
Controller
KEDA scaler
lag-driven
KEDA scaler...
HPA
CPU for stateless APIs
HPA...
Flink autoscaler
reactive parallelism
Flink autoscaler...
Scaled unit
Scaled unit
Ingest API
12 to 90 pods
Ingest API...
Render Service
20 to 160 pods
Render Service...
Channel workers
80 to 600 pods
Channel workers...
Flink task managers
24 to 120
Flink task managers...
Hard ceiling
Hard ceiling
Partition count
concurrency ceiling
Partition count...
Provider TPS
contracted, not elastic
Provider TPS...
Database connections
PgBouncer pool
Database connections...
Backpressure
Backpressure
Shed P2 first
marketing pauses
Shed P2 first...
429 with Retry-After
at the gateway
429 with Retry-After...
Park to DLQ
last resort, replayable
Park to DLQ...
ceiling reached
ceiling reached
vendor cap
vendor cap
still over budget
still over budget
Autoscaling, Capacity and Backpressure
Autoscaling, Capacity and Backpressure
Application we own
Application we own
Security / platform
Security / platform
Queue / topic
Queue / topic
External / third party
External / third party
Data store
Data store
Decision point
Decision point
Interface / broker
Interface / broker
failure / alternate
failure / alternate
Capacity target: 5 000 events/s normal, 25 000 peak, fan-out 1.8, so 45 000 notifications/s peak and roughly 750 M a day. Scaling stops at the partition count, which is why partitions are provisioned for peak and not for today.
Capacity target: 5 000 events/s normal, 25 000 peak, fan-out 1.8, so 45 000 notifications/s peak and roughly 750 M a day. Scaling stops at the partition count, which is why partitions are provisioned for peak and not for today.
v 1.0 · owner Data & AI Global Practice · date 2026-08
v 1.0 · owner Data & AI Global Practice · date 2026-08
Text is not SVG - cannot display

Capacity targets

  • 5 000 events/s normal, 25 000 peak, sustained for 30 minutes
  • Fan-out 1.8, so 45 000 notifications/s at peak and about 750 M a day
  • Maximum event size 64 KB hard, 4 KB at p99 — events carrying more than that are carrying data that belongs behind a reference

Where scaling stops

  • Partition count is the concurrency ceiling, which is why partitions are provisioned for peak rather than for today
  • Provider TPS is contracted, not elastic — beyond it, queueing is the only option
  • Database connections are bounded by PgBouncer, deliberately, so a scale-out event cannot exhaust PostgreSQL

Backpressure order

  • Shed bulk traffic first — marketing pauses before anything transactional degrades
  • Then 429 with Retry-After at the gateway, per tenant, so the producer learns rather than times out
  • Park to the DLQ only as a last resort, because everything parked has to be replayed and deduplicated later