intermediate 3 min answer Multiple choice

At 11:40 a node group scales from 40 to 120 nodes during a traffic spike. The new pods sit in ImagePullBackOff - the public registry is returning 429 for the unauthenticated pulls of base images that three of your images reference directly. Which change removes this failure mode?

container imagesregistryscale-outavailabilitydependencies
Pick one
Show the full answer Hide the answer

The deciding property

A scale-out is the one moment when every layer cache in the fleet is cold. A new node has seen nothing, so it pulls the full image - and at 120 nodes with a 1.2 GB image that is on the order of 96 GB of pulls crossing the internet within a few minutes, from a third party that is entitled to throttle you. The deciding fact is not image size and not the registry's reliability; it is that a dependency on the critical path of capacity acquisition must be inside your failure domain.

Why mirroring at build time wins

Mirroring moves the pull to a registry in the same region, under your availability and your rate limits, and it does so at build time, so the artifact is already local before any node needs it. Cross-region or cross-internet transfer disappears from the scale-out path, pull latency falls from tens of seconds to single digits for a warm regional registry, and the public registry becomes a build-time dependency where a failure costs a delayed release rather than lost capacity.

It also fixes the version of the problem you have not hit yet: a base image that is deleted or retagged upstream. A mirrored copy by digest is immutable and still there.

What it costs

A registry to run or pay for, replication to every region you scale into, garbage collection so storage does not grow without limit, and a build-time step that fails closed when the upstream image is unavailable. Call it 0.1 to 0.2 of an engineer ongoing, plus storage. The decision flips when your fleet is small enough and churns rarely enough that cold pulls are not on any critical path.

Why the other options fail

  • imagePullPolicy IfNotPresent. The layer cache is per node and these nodes are brand new, so there is nothing present. This setting helps a pod restarting on a node that already ran that image, which is not the scenario, and used with mutable tags it quietly pins old bits.
  • Authenticate for a higher limit. This raises the ceiling and keeps a third party in the capacity path. The limit is still finite, still shared across your organisation, and still enforced by someone whose incident review you do not attend. It is a reasonable thing to do and not a fix.
  • Bake images into the node image. Genuinely fast, and it only works for the versions baked at the time the node image was built. The image you need at 11:40 is the one you deployed at 11:00, and rebuilding and rolling node images per release is a heavier pipeline than a mirror. It is a good complement for large stable base layers, not a replacement.
  • Raise the node group maximum. This changes when capacity is requested, not whether the pull succeeds. Earlier scale-out with the same dependency produces the same 429, earlier.

When this is the wrong answer

A three-node single-cluster deployment with one release a week has no cold-pull problem worth a registry. Mirroring there buys an extra component to operate, secure and garbage-collect for a failure mode that never fires. The threshold is roughly: autoscaling in the request path, more than one region, or any node churn you did not schedule.