What happens when a load balancer closes idle connections after 60 seconds while the client's connection pool keeps them for 5 minutes?
Show the full answer Hide the answer
Second by second
At T+60 s the balancer acts on a connection the client still believes is open. The client's pool learns nothing, because an idle socket receives no notification it is programmed to look for.
Two variants follow, and they look nothing alike:
- The balancer sends FIN. The next request written to that socket succeeds locally, the read returns end-of-file, and the client reports a connection reset or an unexpected end of stream. Fast and confusing.
- A device drops the flow silently — a stateful firewall or a managed NAT service, whose TCP idle timeout is commonly 350 seconds in current documentation (2024). Now the write is never answered and the client blocks until its own read timeout. A 5 ms request becomes a 30-second stall, which is the far worse failure.
Where it amplifies
The error rate rises as traffic falls. The chance a pooled connection sits idle past 60 seconds is a function of request rate per connection, so the symptom peaks at 03:00 and in the lowest-traffic region, and it clusters in the minutes right after a peak ends when a large warm pool suddenly goes quiet. That inverted pattern is why the finding is almost always filed as "the network is flaky".
The scale is set by the pool, not by the traffic. At 200 instances holding 50 pooled connections each, a ten-minute quiet period leaves up to 10,000 sockets that are already dead and will fail on first use.
What the user sees, and why a retry may not save them
A sporadic 502 or a one-off multi-second page, roughly one per doomed connection reuse. A retry fixes it only when the request is idempotent and the client treats this specific condition as retryable. Many clients count it as a server error and surface it.
For a non-idempotent write, the silent-drop variant is the dangerous one: the client cannot know whether the request was received. An invisible timeout mismatch becomes duplicate charges or lost writes, which is why this belongs in a design review rather than in a tuning backlog.
What stops it
A rule, not a dashboard. Every idle timeout on the path must be longer than the client's, and the client's must be shortest by a margin: set the client's idle timeout to about 0.8 × the smallest timeout on the path, counting the balancer, every NAT or firewall, and the origin server. Then:
- Enumerate the path's timeouts as configuration, in one place, reviewed when any hop changes. Most occurrences of this are a balancer default nobody read.
- Validate on checkout where the protocol allows a cheap liveness check, so a dead socket is discarded rather than used.
- Honour the server's drain signal. With HTTP/2 a
GOAWAYlets the client stop using a connection before it dies; a client that ignores it will discover the close instead. - TCP keepalive below the NAT idle timeout is a backstop only. It keeps the socket alive in the network and does nothing about a proxy that has ended its own session.
What would have to be true for it to self-heal
The client would need to treat end-of-file on an idle-reused connection, before any response bytes arrive, as a retryable non-error; retry once on a fresh connection; and lower its own idle timeout from the observed age of reset connections. Some clients do all three. Most record a 5xx and emit a metric that blames the server.
When not to add the probing machinery
On a short, trusted path with no NAT and a server you own, derive both ends from one configuration value and skip the liveness checks entirely. The probing machinery earns its place when the path crosses equipment you do not control.