Kingman's Formula
also called VUT Equation, Variability Queueing Approximation
The approximation that splits queue wait into a utilisation term and a variability term, showing that reducing the spread of service times can cut latency as effectively as buying capacity.
Two services run at 80% utilisation. One has a p99 queue wait of 12 ms and the other 400 ms. Nothing in a utilisation dashboard explains the difference, and the usual recommendation — add capacity — is the expensive way to fix only one of them.
The missing variable is variability. John Kingman's 1961 paper The single server queue in heavy traffic approximates the mean queue wait for general arrival and service processes:
Wq ≈ ρ/(1−ρ) × (ca² + cs²)/2 × τ
ρ is utilisation, τ the mean service time, ca the coefficient of variation of inter-arrival times and cs that of service times. The wait is a product of three things, and only one of them is utilisation.
Why it matters
The utilisation term is the famous one: 0.6 → 1.5, 0.8 → 4, 0.9 → 9, 0.95 → 19. The argument for headroom.
The variability term is the useful one, because it is the term you can change without buying hardware. Poisson arrivals and exponential service put both coefficients at 1. Mix 2 ms cache reads with 900 ms report builds in one pool and cs² reaches 4 or more, so the same utilisation produces four times the wait. Synchronised clients — a cron at the top of the minute, retries without jitter — do the same to ca².
The planning consequence: with cs² = 4 at 80% utilisation, halving the service-time spread takes the wait from 10τ to 4τ — about what dropping to 60% would buy — without doubling the fleet.
Implementation patterns
- Separate pools by service-time class. Fast reads and slow aggregations get their own queues, limits and workers. This is the single highest-return application of the formula.
- Jitter every periodic client: retries, polls, cache refreshes, cron jobs. A fixed interval everywhere is an arrival process engineered for maximum variance.
- Pool servers instead of partitioning them. One queue in front of N workers has a much smaller wait than N queues of one, because an idle worker can take anyone's work.
- Timeout the long tail of service time. Capping service time at p99 truncates cs² and improves everyone else's wait, which is why a deadline is a throughput instrument. Report cs² per endpoint so the second term is visible rather than theoretical.
Industry example
A shared training cluster of NVIDIA GPUs is the clearest case. Job durations range from a 20-minute evaluation to a nine-day training run, so cs² is enormous and queue waits become unpredictable at utilisations far below the 85–90% that finance asks for. The fixes operators reach for are separate queues per duration class, preemption for long jobs and reservations — all attacks on the variability term rather than on utilisation. CI farms and batch ETL have the same shape.
Failure scenarios
- A capacity target set from utilisation alone, which passes review and misses its latency objective because variability was never measured.
- One slow endpoint added to a shared pool, which degrades every other endpoint with no change to their own code or traffic.
- Chasing the formula into precision. It is a heavy-traffic approximation for one server, so at low utilisation and with many servers it can be badly off; it gives a direction, not a prediction.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Buy headroom | No code change; helps every endpoint | Fleet cost, for a modest term reduction |
| Split pools by class | Large latency gain at the same cost | More pools to size and alert on |
| Pace admission | Smoother arrivals; predictable queues | A deliberate queue and rejected bursts |
Separating pools fragments capacity: two pools of 30 absorb a burst worse than one pool of 60. So separate by service-time class, where the variance reduction outweighs the lost pooling, and not by tenant or feature, where it usually does not.
When not to use it
If service times are nearly constant and arrivals are scheduled, the variability term is near zero and utilisation can be pushed high. A cron-triggered pipeline over uniform records runs at 90% with trivial queueing, and reserving 40% idle there is waste.
It is also the wrong tool once a queue is bounded and shedding: the wait is capped by queue depth and the cost has moved into the rejection rate.
Interview question
Q: "Two of our services run at 80% CPU. One meets a 50 ms p99 and the other misses a 300 ms p99. Both are the same code base and the same instance type. What do you measure, and what would you change first?"
What a strong answer covers: the service-time distribution rather than the mean; the arrival pattern, looking for synchronised clients and un-jittered retries; utilisation as one of three multiplied terms; splitting pools by service-time class first because it costs no capacity; and that adding instances fixes the symptom at the highest price.
Quick check
Quiz: At 80% utilisation, what single change buys as much latency improvement as dropping to 60%? Halving the coefficient of variation of service time — typically by separating fast and slow work into different pools.
Flashcard: Two services at identical utilisation have queue waits 30× apart. What explains it? — Variability. Kingman's formula multiplies ρ/(1−ρ) by (ca²+cs²)/2, and the second term is the one a mixed workload inflates.