advanced 3 min answer

A capacity test on a live-streaming platform at Twitch scale plateaus at 18000 requests per second however many application instances are added. Fleet CPU is 45%, the database is at 20%, and no queue is growing. What do you check before reporting that ceiling?

load-testingclosed-loopload-generatorephemeral-portsmeasurement
Show the full answer Hide the answer

The first three things I would look at

The load generator, in this order: its arithmetic, its hosts, and its sockets. A plateau that does not move when you add capacity, with every server-side resource idle, is far more often a measurement ceiling than a system ceiling.

  1. Intended rate against issued rate. A closed-loop generator with a fixed number of virtual users cannot exceed users divided by response time. With 1,800 virtual users and a 100 ms response time the generator can issue at most 18,000 requests per second — exactly the number in the report. That coincidence is the whole diagnosis, and it is arithmetic, not telemetry.
  2. Generator host saturation. One pegged core is enough. TLS handshakes, JSON serialisation, and the synchronous write of every result sample are the usual culprits, and a runtime pause in the generator shows up as a gap in sends that nobody is graphing.
  3. Sockets on the generator. Against one destination address and port, a source host has only its ephemeral range to work with — the Linux default of 32768–60999 gives roughly 28,000 ports — and sockets in TIME_WAIT hold theirs. Connection tracking tables on the generator's own path fill for the same reason.

The misleading signal

"No queue is growing" reads as proof the system is healthy, and it is actually the signature of a closed-loop test. In a closed-loop harness each virtual user waits for its response before sending again, so the arrival rate throttles itself to whatever the system can serve. A queue cannot form, by construction. An open-loop generator that issues on a schedule regardless of responses would have produced a growing queue and an honest answer.

The same property corrupts the latency numbers: samples that a real arrival process would have taken during the stall are never taken at all, so the reported p99 describes a traffic pattern no user produces.

How to prove the generator is not the bottleneck

  • Double the generator hosts, hold per-host load constant. If total throughput rises, the generator was the limit. This is the single cheapest experiment and it settles the argument in one run.
  • Calibrate against a trivial endpoint. Point the same harness at a static 200-response on the same network path and record the rate it can issue. That is the harness ceiling, and a test is only credible while it runs below roughly half of it.
  • Record issued-versus-intended rate as a first-class result and fail the run when the gap exceeds about 2%. A load test that cannot state its own error is not a measurement.
  • Run open-loop for the headline numbers, with virtual users sized from the target rate times the target latency rather than from a round number someone typed.

When this is the wrong diagnosis

If generator hosts sit at 20% CPU, issued rate tracks intended rate, and latency climbs as load rises, then 18,000 is the system's number and the search moves to resources that do not appear as fleet CPU: a single-threaded component in the path, a lock or hot row, a NAT device's flow table, a load balancer's per-target connection limit, or a shared network path. The order of that search is set by which resource shows rising latency with flat utilisation.

What a strong answer adds

The organisational layer. A capacity number that was produced by the harness rather than the system becomes a planning input, and six months later a fleet is sized against it. The fix is a standing rule that every capacity claim carries the harness calibration alongside it, in the same document.