Queueing Theory
Why latency explodes as utilisation approaches capacity.
4 to work through
-
intermediate
A capacity review shows services running at 85% CPU. Finance suggests raising it to 95% to save money. What is your response?
2 min answer -
advanced Multiple choice
A platform runs its worker fleet at 85% average utilisation to control cost. Queue times are becoming unpredictable. What does queueing theory say is happening?
2 min answer -
advanced
An inference service's utilisation rises from 70% to 90% and latency more than triples. Why is the relationship non-linear, and what does that imply for capacity planning?
2 min answer -
advanced
WeChat's DAGOR overload control, published at SoCC 2018, profiles a server's load from the average waiting time of requests in its pending queue rather than from CPU utilisation, sheds by business and user priority, and propagates admission levels between services. Why is queuing time the right signal, and why does the propagation matter more than the shedding?
3 min answer
4 terms in this topic
Adaptive Admission Level
A load-dependent priority threshold that a service publishes back to its callers, so requests that would be rejected downstream are never sent and up…
conceptQueueing Delay
The time a request spends waiting rather than being served, which rises non-linearly as utilisation approaches capacity.
conceptUtilisation and Queueing Delay
The non-linear relationship by which waiting time grows as utilisation approaches one, explaining why systems degrade suddenly rather than gradually.
practiceUtilisation Target
The operating point chosen from a latency requirement rather than from cost efficiency, because queueing delay rises non-linearly as utilisation appr…
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.