beginner 3 min answer

A ticket drop at 10:00 takes a site from 200 requests per second to 25000 in under 30 seconds. The application fleet was pre-scaled the night before and sits at 4% CPU. For the first two minutes clients get connection timeouts and 503s that never reach a backend. What happens second by second and what stops it?

cloud-load-balancingpeak-event-readinessdnspre-warmingflash-sale
Show the full answer Hide the answer

Second by second

A managed load balancer is not a box, it is a horizontally scaled fleet that scales the same way your service does, with its own provisioning loop. You are given a DNS name rather than an address, and the provider adjusts the number of balancer nodes behind it in response to observed load. AWS publishes the ELB DNS record with a 60-second TTL exactly so new node addresses can be picked up quickly, and its load-testing guidance has long been to raise load by no more than 50% every 5 minutes. A 125× step in 30 seconds is roughly two orders of magnitude outside that envelope.

At 10:00:00 the existing nodes accept connections to their limit; the rest queue in the accept backlog and then time out. At 10:00:20 the balancer's own saturation metric crosses a threshold and new nodes begin to come up, which takes minutes. As they appear their addresses enter DNS, and clients that resolved at 09:59 keep using the cached addresses for up to a minute. Any client library that resolves once at process start keeps using them indefinitely. Backends stay at 4% CPU throughout, which is the misleading signal: the dashboard everyone opens says there is spare capacity, so the first twenty minutes go into the application.

Where it amplifies

Client retries. Each connection timeout becomes two or three more attempts, and because they resolve from the same cached answer they land on the same saturated nodes. Mobile clients with fixed retry intervals are worse than browsers. Offered load rises in proportion to the failure rate, which is the signature of a retry storm rather than a capacity shortfall.

What stops it

  1. Reserve balancer capacity before the event. Since November 2024 Application and Network Load Balancers support Load Balancer Capacity Unit reservation: you set a minimum capacity and pay for it whether traffic arrives or not. Before that the only route was asking AWS Support to pre-warm the balancer, with a lead time, an expected request rate and a typical response size.
  2. Shape the arrival rather than absorbing it. A queue or waiting room in front of the drop releases a bounded number of users per second, which converts a step into a ramp the scaling loop can follow.
  3. Make a cold balancer cheap to use. Connection reuse through keep-alive, a short connect timeout with jitter, and a retry budget capped as a fraction of requests rather than a per-request retry count.
  4. Confirm the clients re-resolve. A 60-second TTL buys nothing against a library that caches a resolved address for the process lifetime, and this is usually a one-line fix that nobody owns.

When not to pay for reserved capacity

Reserved capacity is paid continuously to protect against a rise that happens rarely. If traffic ramps over tens of minutes it is already inside the autoscaling envelope and the reservation is pure waste. Reserve when the step is scheduled and steeper than about 50% in five minutes — a drop, a broadcast moment, a scheduled market open. Otherwise alarm on the balancer's own saturation metric, not on backend CPU, and leave it alone.