Load Balancer Pre-Warming
also called LCU Reservation, Balancer Warm-Up
Reserving or requesting minimum load-balancer capacity before a scheduled step change in traffic - because a managed balancer is itself an autoscaled fleet whose provisioning loop is slower than a product launch and will return errors while the backends sit idle.
A ticket drop at 10:00 takes a site from 200 requests per second to 25,000 in under 30 seconds. The application fleet was pre-scaled the night before and sits at 4% CPU. For two minutes clients get connection timeouts and 503s that never reach a backend, and the team spends twenty minutes looking at the application because the dashboard says there is spare capacity.
The load balancer is not a box, it is a horizontally scaled fleet with its own provisioning loop. You are handed a DNS name rather than an address, and the provider adds balancer nodes in response to observed load over minutes. Pre-warming puts that capacity in place before the step, and it is a named practice rather than a setting because for a decade the only way to do it was to ask the provider.
Why it matters
Two published numbers bound the problem. AWS publishes the ELB DNS record with a 60-second TTL, so new node addresses take up to a minute to reach clients that re-resolve and never reach a client library that resolves once at process start. And AWS's long-standing load-testing guidance is to increase load by no more than 50% every 5 minutes, a direct statement of how fast the loop moves. A 125× step in 30 seconds is roughly two orders of magnitude outside that envelope.
The failure costs more than its duration suggests, because it lands on the most valuable traffic of the year and because it is misdiagnosed. Backend CPU, the metric everyone watches, is the one metric that proves the problem is not where they are looking.
Implementation patterns
- Reserve capacity where the platform supports it. Since November 2024 AWS Application and Network Load Balancers support Load Balancer Capacity Unit reservation: set a minimum capacity, pay for it whether or not the traffic arrives, and keep autoscaling above it. The older route, still needed elsewhere, is a support request with a lead time, an expected request rate and a typical response size.
- Shape the arrival instead of absorbing it. A waiting room releasing a bounded number of users per second converts a step into a ramp the loop can follow, and is usually cheaper than reserving for peak.
- Warm the targets too. A ramp the balancer survives still lands on cold instances, so readiness gated on filled connection pools plus warm pools or pre-baked images is the companion practice.
- Verify the clients re-resolve DNS, because a 60-second TTL buys nothing against a library that caches a resolved address for the process lifetime.
- Alarm on the balancer's own saturation metric, not backend utilisation, so the next event is diagnosed in the first minute rather than the twentieth.
Industry example
The practice appears in public cloud documentation because the failure is common at scheduled commercial events. AWS guidance has for years named the two cases that need it — expected flash traffic, and a load test that cannot be configured to ramp gradually — and specified what the provider needs to know: the start and end of the window, the expected requests per second, and the typical request and response size. The 2024 reservation feature is the same practice turned into an API, which matters because a capability that requires a support ticket is one nobody exercises during a rehearsal.
Failure scenarios
- Reserving the balancer and forgetting the targets, so the step arrives at a balancer with capacity and a fleet of cold instances, and the tail is just as bad for a different reason.
- Reserving too late, with the reservation still provisioning when the event starts.
- Client-side DNS pinning, where the balancer scales correctly and a population of clients keeps using the original addresses.
- Retry amplification, where each connection timeout becomes two or three more attempts from the same cached address, so offered load rises in proportion to the failure rate.
- Paying the reservation forever, because nobody removed it after the launch and it now sets a floor under a bill that should have returned to baseline.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Reserve minimum capacity | The step is absorbed; no support dependency; rehearsable | Paid continuously whether traffic arrives or not; needs a lead time and a removal step |
| Rely on autoscaling | No standing cost; nothing to remember | Errors for the first minutes of any step steeper than roughly 50% in 5 minutes |
| Shape the arrival with a queue | Cheapest; also protects every downstream component | A product change; users see a waiting room rather than the page |
When not to use it
Do not reserve for a ramp. Traffic rising over tens of minutes is already inside the autoscaling envelope, so a reservation is a continuous payment against a risk that does not exist and the correct control is an alarm on the balancer's saturation metric. Reserve when the step is scheduled and steeper than roughly 50% in five minutes — a drop, a broadcast moment, a market open, a marketing send with a known delivery time.
Do not reach for it when the step is unscheduled either. An unpredictable spike cannot be pre-warmed by definition, and the controls there are admission control, a queue, graceful degradation and static capacity headroom. Pre-warming is for events whose time you know, and substituting it for load shedding leaves the system defenceless against the spikes you did not plan.
Interview question
Q: Your company is running its biggest scheduled product drop next month. Walk me through what you would pre-provision, in what order, and how you would prove the day before that it will hold.
What a strong answer covers: name every autoscaled layer, not just compute — the balancer, the NAT path, the managed database's connection limits, the queue's partition count, any third-party rate limit — and pre-provision each with its own lead time. Reserve balancer capacity, pre-scale and pre-warm the fleet, raise quotas, confirm client DNS behaviour. Then prove it with a load test that reproduces the step rather than a gradual ramp, because a ramp tests the system you will not have. Add the exit: a written step to release the reservation afterwards, and a saturation number from the test to compare against the event.
Quick check
Quiz: Backends are at 4% CPU and clients are getting connection timeouts during a traffic step. What are you looking at, and which metric would have told you in the first minute? Answer: the load balancer's own capacity, which scales on a provisioning loop measured in minutes; the signal is the balancer's saturation or rejected-connection metric, not backend utilisation.
Flashcard: What two published numbers bound a managed load balancer's ability to absorb a step? — A 60-second DNS TTL on the balancer's record, and AWS guidance to raise load by no more than 50% every 5 minutes, which together say anything steeper needs reserved capacity or a queue in front.