LLM Rate Limiting & Traffic Management Service  ·  View 15 of 24  ·  Runtime

Priority, Tiers and Congestion Behaviour

What each tier actually experiences as load rises and as coordination degrades — the two axes read together.

Editable source SVG draw.io All views
Normal · under 70%
Normal · under 70%
Elevated · 70–90%
Elevated · 70–90%
Saturated · over 90%
Saturated · over 90%
Coordination degraded
Coordination degraded
Enterprise · HIGH
Enterprise · HIGH
Serve immediately
Serve immediately
Serve
20% headroom reserved
Serve...
Serve from reserved headroom
Serve from reserved headroom
FAIL OPEN
local bucket × 1.0
FAIL OPEN...
Pro · MEDIUM
Pro · MEDIUM
Serve immediately
Serve immediately
Serve
Serve
Queue 200 ms, then 429
Queue 200 ms, then 429
LOCAL fallback
bucket × 0.8
LOCAL fallback...
Free · LOW
Free · LOW
Serve immediately
Serve immediately
Shed on burst
Shed on burst
Shed first — 429
Shed first — 429
FAIL CLOSED
FAIL CLOSED
Batch · BULK
Batch · BULK
Spare capacity only
Spare capacity only
Deferred to queue
Deferred to queue
Paused
Paused
Paused
Paused
Priority, Tiers and Congestion Behaviour
Priority, Tiers and Congestion Behaviour
Headroom is reserved, not borrowed: the 20% HIGH reserve is deducted from the shared pool at policy publish time, so saturation cannot consume it.
Headroom is reserved, not borrowed: the 20% HIGH reserve is deducted from the shared pool at policy publish time, so saturation cannot consume it.
v 1.0 · owner Data & AI Global Practice
v 1.0 · owner Data & AI Global Practice
Text is not SVG - cannot display

Decisions

  • Headroom is reserved, not borrowed. The 20% HIGH reserve is deducted from the shared pool when the policy is published, so saturation cannot consume it. Borrowing schemes fail exactly when they are needed.
  • Shedding is by tier and is deterministic: BULK pauses, LOW sheds, MEDIUM queues briefly, HIGH is served. A tenant can be told in advance what congestion will feel like.
  • The fourth column is the important one. Congestion and coordination failure are different events, and a design that conflates them behaves unpredictably during an incident.

Numbers

  • MEDIUM queue is bounded at 200 ms and 500 requests per pod, then 429. An unbounded queue converts a rate-limit problem into a latency problem.
  • LOW tier sheds first and receives retry_after derived from the tenant's own window, typically under 2 s.
  • The HIGH reserve is 20% of the shared pool, tunable per deployment; at 100k/s that is 20k/s held back.

Scope note

  • FR12 is marked optional-for-V2 in the brief. It is included in V1 here because the degraded-coordination column has to be answered anyway, and once tiers exist for that, using them for congestion is nearly free.
  • Weighted fair queuing across tenants within a tier is deferred to V2; V1 is first-come within a tier.
  • Cells describe admission behaviour only. What the tenant sees on the wire is in view 11's reason codes.