advanced 3 min answer Multiple choice

A service holds a concurrency limit of 60 with a 50 ms mean service time and serves clients that give up after 2 seconds. Offered load triples. Where should the excess requests wait?

concurrencybounded-queuedeadlinesgoodputload-shedding
Pick one
Show the full answer Hide the answer

The deciding property

The client's deadline sets the queue depth, and nothing else does. Memory does not, because a queue the client has already abandoned holds work that will be thrown away after it is served.

Sixty slots at 50 ms each complete 1,200 requests per second. A request entering at depth N therefore waits about N ÷ 1200 seconds. For that wait to stay inside a 2-second deadline, N must not exceed about 2,400 — and that is the hard ceiling, not the choice. At depth 2,400 a request arrives at the front with its entire budget spent and no time left for the work itself, so the usable depth comes from the latency objective rather than the deadline: for a 400 ms p99 target, 1,200 × 0.4 ≈ 480 slots.

A queue holding roughly one second of work absorbs a short 2× burst at no cost in rejections, which is why "reject everything" is wrong.

What the depth bound buys

Beyond the deadline-derived depth, every additional slot stores a request that will expire before it is served. The service keeps working at 1,200 per second with CPU fully busy while the proportion of completions anyone still wants falls — throughput holds and goodput collapses. That gap is invisible on a CPU graph and obvious on a graph of completions whose deadline had not already passed.

Two additions make the bound work:

  • Check the deadline at dequeue and drop expired entries unserved. Served-but-expired work is the purest waste a system produces.
  • Serve the newest first under overload. FIFO hands the server the oldest and most likely expired request; LIFO or deadline ordering keeps some requests inside budget rather than failing all of them equally.

What would flip the decision, and when not to bound

If this changes Choose Because
Clients have no deadline and the work is durable A persistent queue with no depth bound The work is a job, not a request; latency is not the contract
The service does no CPU work and only proxies The accept backlog and connection limits There is no application queue to bound and the limit is sockets
Protection needed is fairness between tenants A per-tenant quota in front of the queue A single shared queue lets one tenant occupy all of it
Service time varies by orders of magnitude per endpoint A separate limit and queue per class One depth cannot serve 2 ms and 4 s work correctly

Why the other options fail

  • The kernel accept backlog does apply backpressure, and it is the right mechanism for a pure proxy. For a service it is the wrong place: you cannot inspect a deadline, cannot prioritise, cannot report depth, and the client sees TCP retransmission delays rather than a prompt rejection. The queue exists either way — you have only moved it somewhere with no controls.
  • An unbounded application queue is the default in most frameworks and the most common production mistake. It converts an overload that would have been visible as rejections into unbounded latency plus eventual memory exhaustion, and under a growing queue every response is produced after its deadline, so the service works flat out and succeeds at nothing.
  • Raising the limit to 180 assumes concurrency is free. It pushes the service past the point where contention costs more than the added parallelism, so service time rises and the effective completion rate can fall. It also removes the cap that was protecting whatever sits downstream.
  • Rejecting immediately with no queue is defensible and over-tuned. It gives the fastest possible signal and forfeits the smoothing a short queue provides for ordinary burstiness, so a brief 2× burst that the service could have absorbed becomes user-visible errors.