concept

Concurrency

How many operations are in flight at once — the quantity that actually saturates a system, and the one most worth limiting explicitly.

concurrencypoolsbulkheadssaturationlimits

Definition

Concurrency is the number of operations in progress simultaneously. It is distinct from throughput (completed per unit time) and from parallelism (executing at the same instant), and it is the quantity that determines whether a system falls over.

Why concurrency is the right thing to limit

Rate limiting bounds requests per second. But what exhausts a system is in-flight work, and a slow downstream raises concurrency without raising the request rate at all. By Little's Law, concurrency equals arrival rate times latency: at constant traffic, a dependency going from 50 ms to 3 s multiplies concurrency by 60, and a rate limiter notices nothing.

This is why concurrency limiting is more protective than rate limiting and is significantly under-used. A bounded concurrency limit turns a latency problem into a queueing or shedding decision rather than into resource exhaustion.

The pools that matter

Almost every production incident involves one of these being exhausted:

  • Thread or worker pool — how many requests can be processed at once.
  • Connection pool per downstream — and this must be per downstream, not shared, or one slow dependency exhausts the pool for all of them. That is the bulkhead, and it is the single highest-value resilience mechanism and the most frequently missing.
  • Database connection pool — usually the scarcest and least elastic resource in the system.
  • Memory, which bounds how much can be in flight regardless of the other limits.

Setting the limits

Sized to what the resource can actually absorb, not to what feels generous. A database supporting 200 connections and 20 application instances each with a pool of 50 has 1,000 potential connections against a limit of 200 — the failure is guaranteed and will appear under load.

Adaptive concurrency limits — adjusting the limit based on observed latency, in the same way TCP congestion control adjusts its window — are increasingly the right answer, because a static limit is wrong at some point in the system's life.

Failure scenarios

  • A shared connection pool across dependencies, so one slow service exhausts everything.
  • Unbounded concurrency, so a spike consumes all memory.
  • Pool limits summing beyond the downstream's capacity across instances.
  • No queue bound behind the limit, so requests wait indefinitely rather than failing fast.
  • A limit set once, never revisited as the fleet grew.

Interview question

"Why is limiting concurrency more effective than limiting request rate when protecting a service?"