intermediate 2 min answer

At 11:40 autoscaling adds 20 instances to a 40-instance fleet as request rate doubles. By 11:43 p99 has gone from 300 ms to 7 s with a 4% error rate while average CPU across the fleet is 35%. By 11:55 nothing has been changed and p99 is back to 320 ms. What failed and which configuration decision made it possible?

cloud-load-balancingautoscalingslow-starttail-latencyreadiness
Show the full answer Hide the answer

The trigger

The scale-out itself. A newly registered target passes its health check as soon as the process answers, which happens long before it can serve at the fleet's normal latency. An instance that has just booted has an empty local cache, no established connections to the database or downstream services, a cold page cache, and in a JVM or .NET fleet a first-pass interpreter rather than compiled code. It is healthy and it is ten to fifty times slower than its peers for the first thirty to ninety seconds.

Why it propagated

The routing algorithm decided how much damage a cold target could do. Under round robin each of the 20 new targets takes its full 1/60 share from the instant it registers, so a third of all traffic lands on instances that cannot serve it inside the client timeout. Under least outstanding requests the shape is different and not better: a target that has just registered has zero outstanding requests, making it the most attractive target in the group, so it receives a burst and then backs off as its outstanding count climbs.

Average CPU at 35% is the mean of 40 busy instances and 20 instances blocked on I/O, which is why the aggregate dashboard shows nothing. The incident resolving itself at 11:55 with no action is the confirming evidence: the tail was bounded by warm-up time, not by load.

The structural fix, in order

  1. Gate registration on warmth, not liveness. The readiness probe should pass only after the instance has filled its connection pools and served a synthetic request inside the normal latency band. This works on every platform and needs no load-balancer feature.
  2. Ramp traffic into new targets. ALB target groups support slow start, a linear ramp over a configured 30 to 900 seconds. The documented constraint is the interesting part: slow start cannot be enabled together with the least-outstanding-requests or weighted-random algorithms, so you choose between a cold-target ramp and a slow-target-aware algorithm rather than having both.
  3. Remove the cold cost instead of hiding it. Warm pools or pre-baked images move boot and warm-up before the instance is needed, which is the only option that survives a spike faster than the warm-up.

The tempting local fixes both make it worse. Raising the autoscaling threshold delays capacity until the warm fleet is already saturated. Lengthening the client timeout converts a 7-second p99 with errors into a 7-second p99 without them, which is the same user experience with a quieter graph.

When this is the wrong answer

For a stateless service with no local cache that establishes connections lazily in under a second — a Go or Rust service in front of an in-region cache is the usual example — cold start is a non-event, and a slow start ramp only delays capacity you needed three minutes ago. Measure the warm-up curve before configuring a ramp: if an instance reaches normal p99 inside one health-check interval, this mechanism is not your problem and the 7-second tail is somewhere else.