advanced 3 min answer

An admission policy requires that every image have a vulnerability scan with no critical findings, and the admission controller calls the scanner's API at admission time to check. The control works. What has the team bought, what are they paying, and when does the bill arrive?

policy-as-codeadmission-controlattestationcouplingrecovery
Show the full answer Hide the answer

What is gained

A rule that a non-compliant workload cannot exist, rather than one that reports after the fact. That difference is real and worth defending: a pipeline check is bypassable by anything that creates a workload outside the pipeline, and a nightly report tells you what has already been running for hours.

Whether the rule belongs at admission is settled. The question here is what its input costs.

What is paid

Every pod creation now depends on the scanner's availability and latency. Admission webhooks sit in the API server's critical path, and Kubernetes bounds a webhook's timeoutSeconds between 1 and 30 with a default of 10, so the scanner's tail latency becomes the API server's tail latency on every create.

The volume is the part teams underestimate. A few thousand nodes with rolling deploys, autoscaling, CronJobs and evictions produce pod creations on the order of 10,000 an hour in normal operation. A scanner built as a reporting service, called a few thousand times a day by the pipeline, has become a tier-zero dependency called millions of times a month.

Two further costs: the policy is no longer a pure function, so it cannot be tested without a live scanner; and the answer can change between two identical requests, which makes failures non-reproducible.

When the cost becomes visible

During recovery. The moment that matters is a node-pool loss or a zone failure, when the scheduler tries to recreate thousands of pods at once. That is precisely when the scanner is also degraded, because it shares the cluster, and the admission path then decides whether the cluster is allowed to heal itself.

The failure is bimodal and both halves are bad: configured to fail closed, the policy blocks recovery; configured to fail open, the control silently stops existing during the one window an attacker would choose. Neither setting is wrong. The mistake was earlier, in making the decision depend on a network call at all.

How to keep the control without the coupling

Move the fact in-band. The scanner's verdict becomes a signed attestation bound to the image digest, produced once in the pipeline and stored alongside the image. Admission verifies a signature against a public key it already holds, offline, in microseconds.

The control is unchanged: no attestation, no pod. The dependency is gone, the decision is reproducible, and the policy can be unit-tested.

Decision rule: an admission policy must be decidable from the request plus data the controller already holds. If it needs a remote fact, turn that fact into an artefact and attach it to the digest.

The price of the attestation route is key management and expiry handling. An attestation says the image was clean when scanned, so the policy needs a freshness bound, and a critical vulnerability disclosed today invalidates yesterday's attestation for every running workload. That re-scan-and-re-attest loop is real work, and it is work you would need anyway.

When this is the wrong answer

A cluster with a handful of deploys a day, no autoscaling and no multi-zone recovery story can call the scanner live and be fine. The coupling exists but it is never exercised at volume, and signing infrastructure for forty pod creations a day is not a good use of a quarter.

If signing is out of reach, the honest second best is a pipeline-time gate plus a periodic sweep that reports running workloads with no scan record. It is a weaker control and it does not add a tier-zero dependency.