LLM Rate Limiting & Traffic Management Service  ·  View 24 of 24  ·  Assurance

Scaling to One Million Decisions per Second

Design question 6, answered: what scales linearly, what does not, and where the real ceiling actually sits.

Editable source SVG draw.io All views
Tier 1 · Routing
Tier 1 · Routing
Affinity
Affinity
GeoDNS + Anycast
region by latency
GeoDNS + Anycast...
Maglev Hash
on org_id header
Maglev Hash...
Why it matters
Why it matters
Affinity raises lease hit rate
78% → 92%
Affinity raises lease hit rate...
Tier 2 · Decision pods
Tier 2 · Decision pods
limiterd pools
limiterd pools
Pool A
orgs 0–85
Pool A...
Pool B
orgs 86–170
Pool B...
Pool C
orgs 171–255
Pool C...
Unit economics
Unit economics
1 pod ≈ 18k dec/s
4 vCPU · 2 GiB
1 pod ≈ 18k dec/s...
Tier 3 · Coordination shards
Tier 3 · Coordination shards
Valkey Cluster
Valkey Cluster
Shards 0–7
hash tag {org:id}
Shards 0–7...
Shards 8–15
replica per AZ
Shards 8–15...
Unit economics
Unit economics
1 shard ≈ 90k Lua/s
8% of traffic reaches it
1 shard ≈ 90k Lua/s...
Tier 4 · Hot tenant fan-out
Tier 4 · Hot tenant fan-out
Sub-sharding
Sub-sharding
Split above 5k rps
{org:acme#0..7}
Split above 5k rps...
Merge on cool-down
5 min below threshold
Merge on cool-down...
Cost of the split
Cost of the split
Quota divided N ways
overshoot rises to 3%
Quota divided N ways...
Single org on one slot
the real ceiling, not total rps
Single org on one slot...
sticky by org
sticky by org
lease refill only
lease refill only
hot key detected
hot key detected
mitigated
mitigated
Scaling to One Million Decisions per Second
Scaling to One Million Decisions per Second
Interface / broker
Interface / broker
Decision point
Decision point
Application we own
Application we own
Data store
Data store
Risk / gap
Risk / gap
synchronous
synchronous
event / async
event / async
failure / alternate
failure / alternate
Answer to design question 6. Total throughput scales linearly with pods and shards; the binding constraint is a single organisation's keys landing on one Redis slot, which tier 4 exists to relieve.
Answer to design question 6. Total throughput scales linearly with pods and shards; the binding constraint is a single organisation's keys landing on one Redis slot, which tier 4 exists to relieve.
v 1.0 · owner Data & AI Global Practice
v 1.0 · owner Data & AI Global Practice
Text is not SVG - cannot display

The scaling argument

  • Total throughput scales linearly with pods and shards, because a decision touches exactly one tenant's key space and tenants are independent. Getting from 100k/s to 1M/s is 10× the pods and 4× the shards — arithmetic, not architecture.
  • Routing affinity is what makes it cheap. Hashing on org_id at the edge means a tenant's traffic lands on a small pod set, so leases are reused rather than re-fetched: measured lease hit rate rises from 78% with random routing to 92% with affinity.
  • The binding constraint is not total request rate. It is a single organisation's keys landing on one Redis slot, because that is the price of the atomicity in view 12. Tier 4 exists solely to relieve it.

Unit economics

  • 1 limiterd pod ≈ 18k decisions/s on 4 vCPU and 2 GiB. 1M/s needs roughly 56 pods of headroom-free capacity, provisioned at 90 for burst and rolling updates.
  • 1 Valkey shard ≈ 90k script executions/s, and only 8% of decisions reach a shard. 16 shards therefore cover well beyond 1M/s of decisions.
  • The shard count is driven by hot-tenant distribution and memory, not by aggregate throughput.

The hot-tenant trade

  • Above 5k rps a single organisation is split into 8 deterministic sub-shards, {org:acme#0..7}, and merged back after 5 minutes below the threshold.
  • Splitting divides the quota N ways, so overshoot for that tenant rises from about 0.9% to about 3%. That is the explicit cost of the split and it is recorded on the tenant's policy.
  • Sub-sharding is applied automatically on hot-key detection but is visible and overridable, because an automatic action that changes a tenant's enforcement accuracy should never be invisible to the operator.