intermediate 3 min answer Multiple choice

Your manifests are in Git with `spec.replicas: 6` and the reconciler syncs automatically with self-heal enabled. A platform engineer adds a HorizontalPodAutoscaler for the same Deployment with min 6 and max 40. Traffic doubles at 09:00. What happens over the next ten minutes?

gitopsdriftfield-ownershipserver-side-applyhpa
Pick one
Show the full answer Hide the answer

Second by second

09:00 load doubles. The autoscaler's control loop evaluates every 15 seconds by default, so within a minute it has written spec.replicas: 20 and new pods are starting.

09:01 to 09:03 the reconciler's next pass compares the live object with Git, sees 20 where Git says 6, and applies 6. Fourteen pods are terminated. The autoscaler observes the same high utilisation on six pods and writes 20 again. The oscillation period is the reconcile interval, typically a few minutes, and the amplitude is the difference between the declared floor and the load-driven target.

The mechanism underneath is ownership. Kubernetes records which actor last set each field in metadata.managedFields, and a conflict is only reported to an actor that asks for it — a controller applying with conflicts forced simply takes the field back on every pass. Two actors that both force are two actors that both win, alternately, forever.

Where it amplifies

Each reset is an abrupt loss of 14 of 20 pods: in-flight requests are cut, new pods start with cold caches and empty connection pools, and the pods that remain absorb the shed load, which raises utilisation and makes the autoscaler's next target higher than the last. Downscale stabilisation — 5 minutes by default — smooths the autoscaler's own decisions and does nothing about an external writer, so the system's own damping is bypassed.

What the user sees

A saw-tooth in p99, a burst of connection resets every few minutes, and error-budget burn with no deployment in the timeline. The replica count on the dashboard flapping is the giveaway, and it is only visible if that graph exists at per-minute resolution.

Why the other options fail

  • The autoscaler wins through the scale subresource. Plausible, because the autoscaler does write through /scale. The write still lands on spec.replicas of the same object, which the reconciler claims, so the subresource changes the API path and not the ownership.
  • The reconciler wins and the autoscaler gives up. Controllers do not give up. The autoscaler has no memory of being overruled and recomputes from current utilisation on its next tick, which is what makes the loop stable in its instability.
  • Both settle high because reconcilers only add. This confuses a reconciler with a scaling policy. Self-heal is symmetric by design: any difference from the declaration is corrected in whichever direction closes it.

What stops it

Exactly one owner per field. Either delete spec.replicas from the manifest and let the floor live in the autoscaler's minReplicas, or delete the autoscaler. Telling the reconciler to ignore the difference is half a fix: ignoring the diff while the sync still writes the field leaves the reset in place, so the ignore rule has to suppress the write too.

The general lesson for drift: drift that reappears after every reconcile is an ownership bug, not a discipline problem. Drift reports that list changed fields without naming the actor that set them cannot tell the two apart, which is how a drift dashboard becomes 200 alerts a day that everyone mutes.

When not to attach an autoscaler

When the number is a reviewed capacity decision rather than a response to load: quorum sizes for a stateful set, licence-limited pods, a cost-capped batch tier, or a service whose dependency has a hard connection limit. Declare it in Git, and then do not attach an autoscaler.