beginner 3 min answer

Why does making every request 30% faster often leave requests per second exactly where they were, and what did the optimisation actually buy?

throughputlittles-lawutilisationheadroomcost
Show the full answer Hide the answer

The mechanism

In a system that is keeping up, throughput is set by demand, not by capacity. Users send what they send. If arrivals are 400 requests per second and the service can handle 900, then completions are 400 per second — and they stay 400 per second after you make the code faster, because nobody is waiting to be served faster.

Throughput only rises from an optimisation when the system was the constraint: offered load above capacity, a queue growing, work being rejected or timing out. Then service time was the limit and cutting it raises completions directly.

What the 30% actually bought

Little's Law makes the payoff visible. Concurrency equals arrival rate times time in system, so with arrivals fixed and time in system down 30%, in-flight work falls 30%. Fewer requests resident means lower utilisation on every resource they were holding.

The interesting part is that queueing delay falls by much more than 30%. Utilisation is arrival rate times service time, so a 30% cut takes utilisation from 0.70 to roughly 0.49. For a single queue the waiting time scales with utilisation over one minus utilisation, which moves from about 2.3 to about 0.96 — roughly a 2.4× reduction in time spent waiting, from a 30% reduction in time spent working. That is why the p99 improves far more than the p50 after a real optimisation, and it is the number worth reporting.

The other purchase is headroom: the same fleet now saturates at roughly 1.43× the previous arrival rate. Stated in money, you can serve today's traffic on about 30% fewer instances.

The failure mode around this

Two things go wrong, and both are organisational.

The win is never banked. Nobody removes the instances, so a cost saving shows up as nothing. The optimisation is then judged against a throughput number it was never going to move, and the next proposal is declined.

The headroom is spent silently. Traffic grows into the space the optimisation created, utilisation returns to 0.70, and the latency gain evaporates a quarter later with no change to blame. Headroom that is not defended by a target is consumed.

The decision rule before you optimise

Measure whether the system is demand-bound or capacity-bound, and the test is cheap: is anything queueing? Look at queue depth, pool checkout wait, and the gap between offered load and completed load. No queue means you are demand-bound and the optimisation buys headroom and cost, not throughput — so say so in advance and agree which of the two you intend to collect.

If there is a queue, the optimisation raises throughput only if you shortened the bottleneck. Cutting 30% off a stage that is not the constraint changes the total by nothing, which is Amdahl's point.

When this is the wrong framing

For a batch job, throughput is the whole objective and this distinction disappears: the work is all queued at the start, so service time and completion time move together. The same applies to any pipeline fed by a backlog rather than by users.

Common weak answers

  • "The bottleneck must be elsewhere." Possibly, but the usual answer is that there is no bottleneck at all and the request rate is simply what users send.
  • "The measurement is wrong." Both numbers can be correct. Throughput and latency are different quantities and only couple when the system is saturated.