practice

Connection Draining

also called Deregistration Delay, Graceful Shutdown

Allowing in-flight requests on an instance to complete before it is removed from service, rather than terminating them at cutover.

load-balancingdeploymentavailability

Every rolling deployment, scale-in and instance replacement removes a target that is currently serving requests. Without draining, those requests fail — producing the characteristic pattern of a small error spike on every deployment that teams learn to ignore, and that customers experience as random failures.

Draining does the sequence properly: the target is marked as unhealthy or deregistering so no new requests arrive, existing requests are allowed to complete up to a timeout, and only then is the instance terminated.

Getting it right requires the application to cooperate, and this is where it usually breaks. The process must handle the termination signal, stop accepting new work, finish what is in flight, and exit — rather than dying immediately on the signal, which is the default behaviour of many runtimes and container entrypoints.

Two details that catch people. The drain timeout must exceed the longest normal request, or long requests are still cut; and it must be shorter than the platform's forced termination grace period, or the container is killed mid-drain anyway. The two are configured in different places, frequently by different teams, and mismatches are common.

Beyond HTTP: message consumers, WebSocket connections and background workers are not behind the load balancer at all and need their own shutdown handling — finish the current message, commit the offset, stop polling. This is the part omitted in most service templates.