Idle Connection Race
also called Keep-Alive Race, Idle Pool Reset
The unavoidable window where a server closes a pooled connection just as a client sends a request on it, producing a small error floor that is worst when traffic is lightest.
A payments API has an error floor nobody can explain. Roughly one request in 10,000 fails with a connection reset, or a 502 from the proxy in front of it, and the failures cluster between 02:00 and 05:00 when the system is nearly idle. Every backend metric is clean and three sprints of investigation have changed nothing.
The mechanism is a race that cannot be won. A pooled connection sits idle, the server's idle timeout expires and it sends a FIN. While that FIN is in flight the client takes the same connection out of its pool and writes a request. The server, already closing, answers with a reset, and the client cannot distinguish that from a server that failed.
The traffic correlation is the part that misleads everyone. A busy pool has few idle connections, so the window is rarely open; a quiet pool is almost all idle connections sitting near their timeout, so the error rate peaks when load is lowest. HTTP/1.1 gives a server no way to announce a close before a request exists, which is why RFC 9112 (2022) requires clients to be ready to retry on a new connection.
Why it matters
This is one of the most common sources of a permanent small error rate, and it is routinely blamed on the backend because the backend is where the 502 appears.
It also forces a correctness decision that teams get wrong in both directions. Retrying nothing leaves the floor in place forever. Retrying everything turns a reset on a payment into a possible double charge, because a reset says nothing about whether the server processed the request before closing.
Implementation patterns
- Keep the client's idle timeout strictly below the server's, with margin for round-trip time, timer granularity and clock drift: at least 10 seconds under. Managed layer-7 balancers commonly default near 60 seconds of idle (2026), so a backend keep-alive of 75 seconds with a client pool idle of 50 seconds is a sound arrangement.
- Cap connection lifetime as well as idleness. Retiring connections after 5 to 10 minutes means the client chooses when they end.
- Retry once, narrowly: only on a connection-level error, only when no response bytes arrived, and only when the request is idempotent or carries an idempotency key.
- Prefer HTTP/2 or HTTP/3 on internal hops. A GOAWAY frame names the last stream the server processed, so shutdown is unambiguous rather than inferred from timers.
- Do not probe before use. A liveness check on every borrow costs a round trip per request to avoid an event affecting a fraction of a percent.
Industry example
The documented version is vendor guidance rather than a postmortem. Cloud load balancers publish an idle timeout, and their troubleshooting notes for intermittent 502s tell operators to set the backend's keep-alive timeout above the balancer's idle timeout, so the balancer rather than the backend decides when a pooled connection ends. Teams discover the rule from the other side: a framework default of a few seconds of keep-alive behind a 60-second balancer produces a drip of 502s that no backend log explains.
Failure scenarios
- An error floor that rises overnight and falls under load, inverting every intuition about load-related failure.
- A blind retry on a non-idempotent request, producing the duplicate charge that reconciliation finds days later.
- The reflex fix of disabling keep-alive, which trades a 0.01% error problem for a TCP and TLS handshake on every request.
- Health checks keeping one connection warm, so synthetic monitoring never reproduces it and the fault is declared unreproducible.
- Three layers with three idle timeouts - sidecar, application pool, balancer - and nobody owning the arithmetic.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Client idle below server idle | Removes nearly all of the race at no per-request cost | One more coupled configuration value on every hop |
| Narrow single retry | Hides the residual race from users | Needs idempotency discipline on write paths |
| No keep-alive | The race becomes impossible | One to two extra round trips and a TLS handshake per request |
| HTTP/2 internally | GOAWAY makes shutdown exact | A protocol upgrade per hop and stream-limit tuning |
When not to use it
Do not over-apply the mitigations. Dropping the client's idle timeout to a couple of seconds destroys pooling and reintroduces handshake cost on every request, which is a worse trade than the floor. Do not add retries to a write path with no idempotency key; add the key first. Inside a single rack where round-trip time is well under a millisecond the window is so narrow that a connection age cap is the only control worth configuring.
Decision rule: fix it with timeout ordering first, add a narrow retry second, and change protocol only when the hop is important enough to carry HTTP/2 end to end.
Interview question
Q: A service behind a managed load balancer returns a steady 0.01% of 502s, concentrated in the quietest hours, with no matching errors in application logs. Diagnose it, then tell me what changes if the requests are card payments rather than reads.
What a strong answer covers: the idle-close race, and why a quiet pool raises the rate; that the balancer generates the 502 because the backend reset a connection it believed was reusable; the timeout ordering fix stated in the right direction, backend keep-alive above balancer idle and client idle below it; a connection age cap as the second control; and that payments need an idempotency key before any retry is enabled, because a reset is ambiguous about whether the server already committed the work.
Quick check
Quiz: Why does this error rate rise as traffic falls? A quiet pool is mostly idle connections near their timeout, so a larger share of requests are written onto a connection the server is closing.
Flashcard: A pooled HTTP/1.1 connection is reset the moment you reuse it. What ordering rule prevents it, and what does HTTP/2 do instead? — Keep the client's idle timeout about 10 seconds below the server's; HTTP/2 sends GOAWAY with the last processed stream id, so the client knows exactly what was not handled.