pattern

Base Layer Rebase

also called Run Image Swap, Image Rebasing

Replacing the operating-system layers underneath an already-built container image without re-running the application build, so a base-image fix reaches hundreds of services in minutes rather than waiting on every team's pipeline.

container imagescvepatchingsupply chainfleet

A critical vulnerability lands in a C library at 09:00 and the patched base image exists by 09:40. The fleet is 300 services, each built by its own team's pipeline, so the time to patch is set by the slowest team's release process, not by anything the platform controls. Two weeks later a scan still shows 40 services on the vulnerable base, most without a commit this quarter.

A rebase cuts that dependency. An OCI image is an ordered set of layer digests plus a config blob. If the application layers are built so that they do not depend on the specific base layers underneath them, the platform can rewrite the manifest to point at the patched base layers and keep the application layers byte-identical, producing a new image without compiling anything and without the team's build being reproducible today.

Why it matters

Patch latency across a fleet is dominated by coordination, not work. Rebuilding one image takes minutes; getting 300 teams to do it takes weeks. Rebase moves the operation into one platform job, so fleet-wide base patching goes from a multi-week campaign to an afternoon, and it works on services whose teams were reorganised away.

It also removes a quiet correctness risk: a rebuild picks up whatever dependency resolution produces today, so patching an OS library can ship new application dependencies with it. A rebase changes only the layers you intended to change.

Implementation patterns

  • Keep base and application layers independent. A build-time decision: buildpack-style builds separate a run image from application layers by construction, while Dockerfiles that compile against base libraries or interleave apt-get with application steps cannot be rebased safely.
  • Rebase by digest and record both digests, so the operation is auditable and reversible by rewriting the manifest back.
  • Rebase then redeploy. A rebased image in the registry patches nothing; currency is a property of what is running, so the platform job must trigger a rollout and report the share of pods on the new digest.
  • Smoke-test after rebase, because the application was never compiled against the new base. One request through each service's health and main path is enough to catch the class of breakage that matters.
  • Regenerate signatures and attestations. The manifest changed, so provenance and signature records attached to the old digest no longer apply, and admission control will reject the new image if that step is missed.

Industry example

Cloud Native Buildpacks defines this as a first-class operation: pack rebase swaps the run image underneath an existing application image without executing the build lifecycle, specifically so operators can push a security patch across a fleet of images already in production. The separation that makes it possible is part of the buildpack contract rather than something each team has to remember, which is the real lesson for platform teams writing their own image strategy.

Failure scenarios

  • ABI mismatch. The application was linked against the old base's libraries and the patched run image carries a different version, so the process starts and fails on a code path nobody smoke-tested.
  • Rebased and never rolled out. Scanners go green against the registry while production stays vulnerable.
  • Attestation drift, where signature verification at admission rejects the rebased digest and the rollout stalls.
  • False comfort. The vulnerability is in the application's dependency tree, not the base, and rebase cannot touch it.

Trade-offs

Choose Gains Pays
Rebase as the fleet patch path Hours instead of weeks to close a base CVE; no dependency on team pipelines A build convention every image must follow and a smoke-test harness per service
Rebuild everything Picks up application dependency fixes too Patch latency set by the slowest team, and new application changes shipped during an incident

When not to use it

With 15 images from one well-maintained Dockerfile set and a 10-minute rebuild, a scripted rebuild-and-roll is simpler and catches more. Rebase earns its constraints at roughly 50 or more independently built images, or when teams' pipelines cannot be relied on to build on demand. It is also the wrong tool when the vulnerability is in application dependencies.

Interview question

Q: Your scanner reports 300 services on a vulnerable base image. Leadership wants the fleet clean in 24 hours. What do you actually do, and what could stop it working?

What a strong answer covers: the difference between patching the registry and patching what runs; rebase as a manifest rewrite rather than a build; the prerequisite that application layers not depend on base layers; smoke tests because nothing was recompiled; regenerated signatures so admission does not reject the result; the images that must be excluded; and that application-level vulnerabilities still need a rebuild.

Quick check

Quiz: Why is "all images rebased" not the same as "the fleet is patched"? Because the running pods still hold the old digest until a rollout happens, so currency is measured on the cluster and not in the registry.

Flashcard: What does a base layer rebase change and what can it not fix? It swaps OS layers under existing application layers in seconds with no rebuild, and it cannot fix vulnerabilities in the application's own dependencies or breakage from an ABI change.