A checkout service can serve 100 requests per second and receives 60 per second all day. Why do requests still spend time waiting, and what would have to be true for the wait to be zero?
Show the full answer Hide the answer
The mechanism
Capacity is not storable. A second in which the server is idle does not bank service for the next second, so the only thing that matters is whether a request finds the server busy at the moment it arrives. At 60% utilisation, roughly 60% of arrivals find it busy, and those requests wait.
"100 per second" and "60 per second" are averages over a day. The arrivals are not spaced 16.7 ms apart. In a typical second you might get 48, then 71, then 55. Within each second the gaps are uneven too, and two requests that arrive 2 ms apart cannot both be served first.
The two sources of variability
Kingman's 1961 approximation for a single server names them explicitly. Mean queue wait is roughly
utilisation term × variability term × mean service time = ρ/(1−ρ) × (ca² + cs²)/2 × τ
where ca is the coefficient of variation of the gaps between arrivals and cs that of service times. With random (Poisson) arrivals and exponentially distributed service, both squared coefficients are 1, the variability term is 1, and ρ/(1−ρ) at ρ = 0.6 is 1.5. A service at 60% utilisation makes the average request wait about one and a half service times before it starts.
The utilisation term is the famous part: 0.6 → 1.5, 0.8 → 4, 0.9 → 9, 0.95 → 19. The variability term is the useful part, because it is the one you can change without buying hardware.
What this changes in practice
Zero wait needs zero variability, not spare capacity. If arrivals were perfectly spaced and every request took exactly the same time, both coefficients would be 0 and the wait would be 0 at any utilisation below 1. Real systems get partway there:
- Reduce cs² by splitting fast and slow work into separate pools. One queue serving 2 ms reads and 900 ms report builds has a huge service-time variance, and the reads pay for it.
- Reduce ca² by smoothing arrivals: jitter on client retries and polling, admission at a paced rate, batch windows. Synchronised clients are the most expensive arrival pattern there is.
- Pool servers rather than partitioning them. One queue in front of 4 workers waits far less than 4 queues of 1, because an idle worker can take anyone's request.
- Buy headroom last, because it is the most expensive lever: moving from 60% to 30% utilisation cuts the utilisation term from 1.5 to 0.43, which doubles your machine bill for a 1-service-time gain.
When this is the wrong answer
If arrivals are scheduled and service times are nearly constant, high utilisation is safe. A batch pipeline triggered by cron, processing uniform records, can run at 90% with trivial queueing, and holding 40% idle there is waste. The same applies to a bounded queue with shedding: the wait is capped by the queue depth, and the cost shows up as a rejection rate rather than as latency.
Common weak answers
- "It is at 60%, so there is no queue." Utilisation is an average; queueing is driven by the distribution around it.
- "Add capacity." It works and it is the costly option. Check the service-time spread first; teams routinely find one endpoint contributing most of the variance.
- "Average response time is fine." The average hides the arrivals that landed behind a slow request, which is exactly the population that notices.