beginner 3 min answer Multiple choice

An inference service's container image is 6.2 GB because the model weights sit in the final layer. Autoscaling adds twelve fresh nodes at peak; the new pods stay NotReady for several minutes and other pods scheduled onto the same nodes wait behind them. What is happening and which change helps most?

container imagesimage pullcold startautoscalingnode cache
Pick one
Show the full answer Hide the answer

What happens second by second

A fresh node has an empty image store, so the 6.2 GB must be transferred and unpacked before the container starts. At an effective 1.25 Gbps that is roughly 40 seconds of transfer in the best case, and decompression and writing layers to disk often cost as much again. Twelve nodes pulling the same image at once turn one registry endpoint into the bottleneck and the per-node time grows.

Then the queueing effect that the question is really about: the kubelet pulls images serially by default, one request to the image service at a time, and the Kubernetes documentation notes it never pulls multiple images in parallel for a single pod. So a small sidecar image or another team's pod scheduled onto the same new node waits behind your 6.2 GB pull. One heavy image degrades every cold start on that node.

The load-bearing insight: image size is paid per node and per version, not per release. A 6 GB image costs nothing extra on a warm node and costs minutes on every scale-out, node replacement, spot reclamation and cluster upgrade. Teams size images against release frequency and are then surprised by node churn.

The change that actually helps

Split the artefact by change rate. The application layer changes daily and is small; the weights change rarely and are huge. Ship a 200 MB image and fetch weights at startup from object storage into a node-local cache — a host path or a volume shared by pods on that node — so the second pod on the node starts in seconds. The pull is then proportional to what changed, the registry stops being a scale-out dependency, and base-image patching gets cheaper because there is no multi-gigabyte layer to rebuild and redistribute.

Pair it with two cheap measures: a pull-through registry mirror inside the cluster, and a pre-pull of the current image onto new nodes as part of node bootstrap so the first pod is not the one paying.

Why the other options fail

  • Raise the readiness threshold. This stops the probe from killing a pod that is merely slow, which is sometimes necessary. It does not make capacity arrive sooner, and it hides the problem: cold-start latency stops being an alert and becomes an unmeasured property of every scale-out.
  • Disable serialized pulls. Parallel pulls help a node that needs several different images, and they are a reasonable tuning step. Here the bottleneck is bytes over one network path and CPU for decompression; running more pulls at once divides the same bandwidth and does nothing for the single heavy pull, which is also never parallelised within one pod.
  • Keep spare replicas warm. This is the right answer when the spike is predictable and short, and it is how many teams buy time. It pays for idle accelerators continuously to hide a fixable pull cost, and the moment demand exceeds the warm pool you are back to a multi-minute cold node.

Common weak answers

"Use a smaller base image" is good hygiene and irrelevant at this scale: distroless saves tens of megabytes against 6 GB of weights. "Add a bigger registry" treats a symptom whose cause is that the same bytes are shipped to every node on every release. The decision rule worth remembering: anything in an image that changes at a different rate from the code belongs outside the image, and the threshold where that starts to matter is roughly the point where pull time exceeds your scale-out latency target.