intermediate 3 min answer Multiple choice

At 09:10 a checkout service starts returning 503s. Its autoscaler has taken the fleet from 30 to 180 pods in twenty minutes, CPU per pod sits at 22%, and the database reports connection waits. That day's compute bill is four times normal and traffic is up only 30%. Which explanation fits every symptom?

autoscalingconnection poolingfeedback loopcost spikesaturation
Pick one
Show the full answer Hide the answer

The trigger

Traffic rose 30%, which pushed latency up slightly. The autoscaler reacted. Nothing about the trigger is unusual; the damage comes from what scaling out consumes.

Why it propagated

Each pod holds its own connection pool. At 10 connections per pod, 30 pods hold 300 connections against a server limit of, say, 500. At 180 pods the fleet is asking for 1800. Past the limit, connections are refused or queue, so request latency rises, so the autoscaler's target is still missed, so it adds more pods, each of which claims more connections. That is a positive feedback loop: the remedy is the cause.

CPU at 22% across 180 pods is the clue that settles it. A fleet six times larger than it was, sitting at a fifth of its capacity, is not working harder. It is waiting.

Why the bill moved with the errors

Two effects compound. The pods themselves cost six times more for the hours they ran, and the cluster autoscaler added nodes underneath them, which are billed per hour and scale in far more slowly than pods do because of drain and pod-disruption rules. A fleet that overshoots for twenty minutes typically holds the nodes for hours. A cost spike arriving at the same minute as an error spike is evidence of a scaling loop, not of demand, and that pairing is the cheapest detector you have.

Why the other options fail

  • A larger database instance class. A bigger machine raises the connection ceiling, so the loop runs further before it breaks and the bill grows on both sides. It would be the right answer if connections were within limits and CPU or I/O on the database were saturated, which the symptoms do not show.
  • CPU throttling by the kubelet. Throttling produces high CPU utilisation against the limit and latency that does not improve with more pods. Here CPU is 22%, so there is nothing to throttle. This is a real and common cause of latency under autoscaling, which is why it makes a convincing distractor.
  • A genuine spike that the fleet absorbed correctly. Correct absorption looks like rising CPU and stable latency. Six times the fleet for 1.3 times the traffic, with errors, is not that.

The structural fix

  1. Put a hard maximum on replicas derived from the shared resource, not from a budget guess: maximum pods is roughly the connection limit divided by the per-pod pool size, with headroom left for migrations and maintenance.
  2. Multiplex through a connection pooler in front of the database so pod count and connection count stop being proportional.
  3. Scale on a saturation signal for the real constraint, such as connection acquisition wait time or queue depth, so the controller can tell "waiting" from "busy".
  4. Alert on cost per request per hour, which moves the moment a scaling loop starts and does not move during honest growth.

When this is the wrong answer

If the pool is already multiplexed and CPU genuinely sits near 80% at peak, the fleet growth is correct and the extra spend is the price of the traffic. The reflex to cap replicas is then harmful, because the cap becomes an availability limit during a real event. The rule is narrow: any autoscaler whose scale-out consumes a fixed shared resource needs a maximum derived from that resource. Where scale-out consumes nothing shared, let it run.