pattern

Load Shedding

Deliberately rejecting a portion of incoming work so the system continues serving the remainder at acceptable latency, rather than degrading for everyone.

overloadresiliencecapacity

The alternative to shedding is not "serve everything". It is congestive collapse: queues grow, latency rises past the point of usefulness, callers time out and retry, and the system delivers nothing while working at full capacity. Shedding is how a system stays useful at overload instead of becoming uniformly useless.

The decision to make explicitly is what to shed, and this is a business question dressed as a technical one. Shedding at random is the crudest option. Better: shed by priority, so health checks and payment completion survive while recommendations and analytics do not; shed by cost, dropping expensive queries first; shed by tenant, protecting paying customers or enforcing fairness so one caller cannot consume the pool; and shed requests whose deadline has already expired, which is free capacity nobody is waiting for.

The mechanics that make it work: reject early and cheaply, at the edge or on admission, because a rejection that costs as much as serving the request has not helped. Return a clear signal — a 429 or 503 with a retry hint — so well-behaved clients back off rather than retrying immediately.

The measurement to have in place beforehand: concurrency in flight, not CPU. Queue depth and in-flight count are the signals that predict collapse; utilisation is the one that lags it.