pattern

Hedged Request

also called Request Hedging, Tied Request

Sending the same request to a second replica after a short delay and using whichever responds first, to cut tail latency caused by unlucky slow servers.

tail-latencyresilienceperformance

The insight behind it is that tail latency is usually not caused by a slow request but by a slow server — a node doing garbage collection, sharing a host with a noisy neighbour, or having just lost its cache. The request would have been fast almost anywhere else.

Hedging exploits that. Send the request; if no response arrives by the 95th percentile latency, send it to another replica and take the first answer. Because only 5% of requests trigger a hedge, the extra load is around 5%, and the effect on p99 and p99.9 is often dramatic — this technique is one of the standard tools in large-scale serving systems for exactly this reason.

The conditions that make it safe are specific and must be checked. The operation must be idempotent, since both copies may execute. There must be spare capacity, because hedging under overload adds load precisely when the system is struggling, which turns a latency optimisation into an amplifier. And the trigger must be a percentile computed from live traffic rather than a fixed value that becomes wrong as conditions change.

The refinement worth knowing: tied requests, where the two replicas are told about each other so that whichever starts first cancels the other, cutting the wasted work substantially.