metric

Throughput

Work completed per unit time — bounded by the system's narrowest resource, and traded against latency.

throughputcapacityconcurrencybottleneckscaling

Definition

Throughput is completed operations per unit time. It is bounded by whichever resource saturates first, and increasing anything else has no effect at all — which is why throughput work must start with identifying the bottleneck rather than with optimisation.

The relationship with latency

They are related through concurrency, and confusing them is common:

Throughput = concurrency ÷ latency

So throughput can be increased either by reducing latency or by increasing concurrency. Increasing concurrency beyond the point of saturation does not increase throughput — it increases queueing, so latency rises while throughput stays flat. Past that point it usually decreases throughput as context switching, contention and memory pressure take over.

This produces the characteristic shape: throughput rises with load, plateaus, and then falls. The region past the plateau is where systems collapse, and staying out of it is what admission control and load shedding are for.

Improving throughput, in order

  1. Find the bottleneck. Everything else is guessing. It is usually one of: CPU, memory bandwidth, disk I/O, network, a lock, a connection pool, or a downstream dependency.
  2. Do less work per operation. Caching, avoiding N+1 patterns, cheaper serialisation. The largest wins are usually here and require no additional capacity.
  3. Batch. Amortising per-operation overhead — round trips, transaction setup, syscalls — is frequently an order of magnitude, and it trades a little latency for a lot of throughput.
  4. Increase parallelism, up to the saturation point.
  5. Add capacity, once the above is exhausted.

Batching as the central trade

Batching is the clearest expression of the throughput/latency trade: waiting to accumulate a batch costs latency and buys throughput. The right batch size and timeout follow from the latency budget, and the mistake is choosing a batch size without reference to it.

Failure scenarios

  • Adding capacity behind an unmoved bottleneck — more application servers in front of a saturated database achieve nothing but a larger bill.
  • Unbounded concurrency, pushing the system past the plateau into collapse.
  • Measuring throughput without latency, so a system that is "handling 10,000 requests per second" is doing so at 30 seconds each.
  • Optimising the fast path while the slow path holds the bottleneck.

Interview question

"Throughput has plateaued and adding instances does not help. What are the candidate bottlenecks and how do you identify which?"