concept

Image Pull Amplification

also called Per-Node Image Cost, Cold Node Pull Cost

The property that a container image's size is paid on every node that has never seen that version, so its real cost scales with node churn and scale-out events rather than with release frequency.

container imagesimage pullcold startautoscalingnode churn

An inference service ships a 6.2 GB image because the model weights sit in the final layer. Releases are weekly and nobody minds the build time. Then autoscaling adds twelve fresh nodes at peak, the new pods sit NotReady for several minutes, unrelated pods scheduled onto the same nodes start slowly too, and the platform gets a ticket about capacity that is really a ticket about bytes.

Image size is not paid per release. It is paid per node that does not already hold that version. Teams reason about images through their build and push, which happens once, and the pull happens on every node the image lands on for the first time: every scale-out, every node replacement, every spot reclamation, every rolling cluster upgrade, every eviction that reschedules a pod elsewhere. In a cluster with healthy node churn that multiplier is large and invisible in any dashboard the owning team looks at.

The second amplifier is local: the kubelet pulls images serially by default, one request to the image service at a time, and Kubernetes' documentation notes it never pulls multiple images in parallel on behalf of one pod. So a heavy pull does not only delay its own pod; it queues every other first-time pull on that node behind it.

Why it matters

The arithmetic is unforgiving. 6.2 GB at an effective 1.25 Gbps is roughly 40 seconds of transfer in the best case, and decompression plus writing layers to disk often costs as much again, so a cold start floor of 90 seconds to 3 minutes is normal. Twelve nodes pulling the same image at once make one registry endpoint the bottleneck and each node's time grows. Your autoscaler's reaction time becomes irrelevant when capacity cannot serve for minutes after it arrives, and every capacity plan built on scale-out latency is wrong by that margin.

It also silently raises the cost of security work: a multi-gigabyte layer has to be rebuilt and redistributed to the whole fleet for a patch that changed 4 MB of libraries.

Implementation patterns

  • Split artefacts by change rate. The application layer changes daily and is small; model weights, datasets and large static assets change rarely and are huge. Keep a 200 MB image and fetch the large artefact at startup into a node-local cache shared by pods on that node.
  • Run a pull-through registry mirror inside the cluster so a scale-out event does not become external egress and a rate limit on someone else's service.
  • Pre-pull the current version during node bootstrap, so the first pod on a new node is not the one paying. This is the cheapest fix available and it is often one DaemonSet.
  • Order layers by volatility so the frequently rebuilt layers are the small ones, and keep the base pinned by digest so unchanged layers are genuinely reused.
  • Enable parallel pulls with a bounded limit where nodes routinely need several different images, remembering that bandwidth and decompression are shared and that one pod's images are never parallelised.
  • Measure it: p95 time from pod scheduled to container started, split by whether the node already had the image. Two populations in one histogram is how this stays hidden.

Industry example

The clearest public ground truth is the platform documentation rather than a company blog: Kubernetes documents serial image pulls as the kubelet default and states that multiple images are never pulled in parallel for a single pod, which is the mechanism behind the "unrelated pod is slow too" symptom that teams usually attribute to the scheduler. The pattern shows up in production wherever machine-learning images are built the obvious way — weights copied into the image so the artefact is self-contained — and it is one of the few platform problems whose fix is almost free once named.

Failure scenarios

  • Readiness timeouts on new nodes. The probe fails before the pull finishes, the pod is killed and restarted, and the pull begins again; at multi-gigabyte sizes this can loop.
  • Registry as a single point of failure at peak. The moment you need capacity most is the moment every new node pulls, and a registry outage or rate limit becomes a capacity outage.
  • Disk pressure and eviction churn. Several versions of a 6 GB image on one node fill the image store, garbage collection evicts images that are about to be needed, and pulls repeat.
  • Hidden cost migration. Teams add warm spare replicas to mask the delay, which converts a fixable pull cost into permanent idle accelerator spend.

Trade-offs

Choose Gains Pays
Large self-contained image one artefact to promote and attest; no runtime fetch to fail minutes of cold start per node and expensive fleet-wide patching
Small image plus fetched artefact fast cold starts and cheap patching a second distribution path with its own cache and failure modes

The honest cost of splitting is that the running container is no longer described entirely by its image digest, so provenance and reproducibility need the fetched artefact's digest recorded too.

When not to use it

If images are already small — a few hundred megabytes — and the cluster's node population is stable, the amplification is a few seconds and not worth a second distribution mechanism. Keep the single artefact. The threshold to act is behavioural rather than absolute: when pull time exceeds your scale-out latency target, or when a base-image patch takes longer to distribute than to build, the large artefact belongs outside the image.

Interview question

Q: A team's 6 GB image gives multi-minute cold starts during autoscaling and they ask for a faster registry. What do you propose instead, and what do you measure to prove it worked?

What a strong answer covers: naming per-node-per-version as the cost basis and node churn as the multiplier; the serial-pull mechanism that makes it a node-wide problem; splitting by change rate with a node-local cache; a pull-through mirror and bootstrap pre-pull as cheap immediate wins; rejecting readiness-threshold tuning as hiding the signal; and the measurement — p95 scheduled-to-started split by warm and cold node, plus days between a base patch and fleet coverage.

Quick check

Quiz: Why does a 6 GB image cost nothing on Tuesday and three minutes on Wednesday? — Because the cost is paid per node that has not seen that version, so it appears on scale-out and node replacement rather than on release.

Flashcard: Why does one heavy image slow down unrelated pods on the same new node? — The kubelet pulls images serially by default so other first-time pulls queue behind it.