Container Image Strategy
Base images, layer hygiene, rebuild cadence, and patching a fleet of images.
4 to work through
-
beginner Multiple choice
An inference service's container image is 6.2 GB because the model weights sit in the final layer. Autoscaling adds twelve fresh nodes at peak; the new pods stay NotReady for several minutes and other pods scheduled onto the same nodes wait behind them. What is happening and which change helps most?
3 min answer -
intermediate Multiple choice
A critical CVE in a widely-used C library is announced at 09:00. You have 200 services in containers, each built by its own team's pipeline. Which approach patches the fleet fastest?
3 min answer -
intermediate
A platform must define a container image strategy for many teams. What decisions matter, and what is the security constraint?
2 min answer -
intermediate Multiple choice
At 11:40 a node group scales from 40 to 120 nodes during a traffic spike. The new pods sit in ImagePullBackOff - the public registry is returning 429 for the unauthenticated pulls of base images that three of your images reference directly. Which change removes this failure mode?
3 min answer
4 terms in this topic
Base Image Currency
How far behind the fleet's running images are from their patched base, which is the number that decides how fast a critical vulnerability can be closed.
patternBase Layer Rebase
Replacing the operating-system layers underneath an already-built container image without re-running the application build, so a base-image fix reach…
practiceContainer Image Strategy
The organisational decision about where images come from, what they are built on, and how a base layer fix reaches everything already deployed.
conceptImage Pull Amplification
The property that a container image's size is paid on every node that has never seen that version, so its real cost scales with node churn and scale-…
Neighbouring topics
Platform Engineering
General material on internal platforms as products with users, adoption and lifecycles.
Internal Developer Platform
The assembled surface teams actually touch, and what belongs behind it.
Paved Road & Golden Path
A supported default route that is easier than the alternatives rather than mandatory.
Self-Service Provisioning
Teams getting infrastructure without a ticket, and the guardrails that make that safe.
Service Templates
Scaffolding new services with observability, CI and security already wired in.
Platform APIs
Treating the platform's own interfaces as contracts with consumers and compatibility rules.
Platform Tenancy
Isolating teams sharing a cluster, account or pipeline fleet, and where isolation must be hard.
Cluster Architecture
How many clusters, split by what, and the blast radius each split buys.
Service Mesh Operations
What a mesh genuinely solves, its failure modes, and the cost of running one.
Developer Environments
Local, remote and ephemeral environments, and the fidelity each can honestly claim.
Inner Loop & Outer Loop
Where an engineer's time actually goes, and which loop a platform investment shortens.
Abstraction Level Choice
How much to hide, and the leak that turns a helpful abstraction into a trap.
Platform SLOs
Committing to reliability for internal consumers who cannot choose another provider.
Platform Adoption
Migrating existing teams onto a platform without a mandate, and reading the adoption curve.
Platform Funding
Central cost, showback, chargeback, and justifying a team that ships no customer feature.
Platform API Deprecation
Removing something dozens of internal teams depend on, on a timeline that holds.
Guardrails vs Gates
Preventing a class of mistake automatically versus stopping to ask a human.
Platform Telemetry
Instrumenting the platform itself: usage, friction, and where teams leave the paved road.
Platform Team Topologies
Stream-aligned, enabling, complicated-subsystem and platform teams, and their interactions.