A load-test report claims 5000 requests per second sustained at a mean response time of 20 ms. The harness ran 60 virtual users with no think time. Review the report.
Show the full answer Hide the answer
The arithmetic that settles it
Little's Law, proved by John Little in 1961: L = λW = 5000 × 0.020 = 100 requests in flight on average. A closed-loop harness with 60 virtual users and no think time holds at most 60 in flight, because each user waits for its response before sending again. The report's three numbers cannot all be true.
Run it the other way to get the bound: 60 users at 20 ms each can issue at most 60 ÷ 0.020 = 3000 requests per second. The claimed throughput is 1.7× what the stated concurrency and latency permit.
The three explanations, and how to tell them apart
- Throughput is over-reported. The harness counted attempts, submissions or bytes rather than completed responses. Check: completed-request count ÷ test duration should equal the reported rate.
- Latency is under-reported. Classic coordinated omission — the harness stalls while the system is stalled, so the slow samples a real arrival process would have produced are never recorded. Check: compare the sum of all recorded response times against wall-clock test duration × user count.
- Concurrency was higher than 60. An asynchronous client that fires without waiting makes "virtual users" a thread count rather than a concurrency level. Check: plot in-flight requests over time.
All three make the system look better than it is, and a capacity plan derived from this report under-provisions by about 40%.
What to change in the test
Report in-flight concurrency as a time series next to throughput, so the identity L = λW can be checked on the chart rather than in an argument. Drive the test open-loop from a target arrival rate rather than from a user count, so the generator keeps arriving when the system slows, which is what real traffic does. Reconcile λ from the completion log, not from the generator's intent.
Why the other options fail
- "The result stands." It cannot, and accepting it publishes a capacity number roughly 1.7× too high. This option is the common failure: the numbers are each individually plausible, so nobody multiplies them.
- "More virtual users." The right instinct at the wrong moment. More load is how you find a ceiling, and you cannot interpret a report whose own numbers contradict each other. Raising users in a closed loop also raises concurrency and latency together, which tells you less about arrival-rate behaviour than an open-loop run.
- "Latency is implausible." 20 ms is ordinary for a cached read path. Rejecting a number for being small is a habit that gets the wrong answer here and distracts from an arithmetic error that takes ten seconds to find.
When this is the wrong answer
If the generator is open-loop or asynchronous, "60 virtual users" is a thread count and not a concurrency level, and the report may be correct while being mislabelled. The reviewer's job is then to get the in-flight series rather than to reject the result. Ask which mode the harness ran in before accusing anyone of bad arithmetic; an open-loop generator with 60 threads can easily hold 100 requests in flight.
What a strong reviewer adds
Little's Law is a consistency check on every performance claim you are handed, including your own monitoring, and it costs nothing to apply. In-flight concurrency, throughput and latency are three views of two degrees of freedom, so any report giving all three is checkable, and a surprising number of them fail. Prefer dashboards that show all three for exactly this reason.