Admission Control
Deciding at the entrance whether to accept a request at all, based on whether the system can complete it within its deadline.
Admission control is the discipline of refusing work early rather than accepting everything and degrading for everyone. It rests on a single question the system must always be able to answer: given current capacity and queue depth, can this request be completed before it stops being useful? If not, reject it now.
The counter-intuitive part is that rejecting quickly is more generous than accepting slowly. A client that receives a fast rejection can back off, fail over, degrade its own experience, or tell the user something honest. A client left waiting past its own timeout gets nothing, having consumed resources that could have served someone else.
Why it matters
Without admission control, a system under overload does the worst possible thing: it accepts all work, completes none of it within a useful time, and spends its full capacity producing results nobody is waiting for. Throughput of useful work collapses to near zero while utilisation reads 100%.
Implementation patterns
- Bounded queues with a wait-time estimate. Reject when projected wait exceeds the request's deadline rather than when the queue is full — depth is a poor proxy for time.
- Priority classes. Interactive above batch, paid above free, in-flight sessions above new ones. Shedding uniformly wastes the ability to protect what matters.
- Per-tenant quotas at the edge, so one tenant's burst is rejected before it consumes shared capacity.
- Deadline propagation, so downstream services can also refuse work whose caller has given up.
- Cheapest-possible rejection. The check must happen before authentication does expensive work, before a connection is allocated, and certainly before a scarce resource is reserved.
- Load-based rather than rate-based signals where possible: queue latency and resource saturation adapt automatically as capacity changes, whereas a fixed request rate is wrong the moment the hardware or the workload mix changes.
Industry example
An AI inference platform is the sharpest illustration, because the resource is both scarce and expensive and a single request occupies it for seconds. The ordering is forced by economics: reject at the edge in microseconds, or discover after a GPU has been generating for four seconds that the client timed out three seconds ago.
The full ladder is admission control and rate limits at the edge; then priority queues with bounded wait, so a request that cannot meet its class deadline is refused immediately; then dynamic batching at the accelerator, which is where the actual capacity comes from; with autoscaling as a slow background control that responds to trends and cannot respond to bursts at all.
The same structure appears in database connection admission, in CI build queues after a large release, and in payment processing during peak retail events. The specifics differ; the principle — never let a scarce resource be consumed by work that will be discarded — does not.
Failure scenarios
- Unbounded queues, which convert fast honest rejection into slow expensive rejection and then into memory exhaustion.
- Admission control after the expensive step, which is decoration.
- Uniform shedding, dropping a critical checkout at the same rate as an analytics ping.
- Rejecting without a retry signal. A rejection with no
Retry-Afterand no distinction between "try again shortly" and "you are over quota" causes immediate retries, which is amplification. - Static thresholds tuned once for a fleet whose capacity has since doubled.
Trade-offs
Admission control means deliberately refusing requests the system might have served, and it will sometimes be wrong — a conservative threshold sheds during a spike that would have passed. It also adds a decision to the hot path and a new set of thresholds to tune and keep correct.
What you buy is the ability to remain useful under overload rather than uniformly useless, and the ability to keep promises to the traffic you did accept. For any system with a meaningful SLO, that trade is decisively worth it.
Interview question
"Your service receives requests faster than workers can process them. Talk me through bounded queues, backpressure, admission control, load shedding and prioritisation — and tell me what you would do differently for a checkout request versus an analytics event."