advanced 3 min answer

Review this. An admission policy blocks any image lacking a signed provenance attestation, verified against a key held in the cluster's secret store. 94% of deployments pass on the first attempt, no attestation has been rejected in two quarters, and teams may register self-hosted build runners with the same build system. What would you change, what would you remove, and what would you leave alone?

provenanceadmission-controlslsabuilder-identitytrust-root
Show the full answer Hide the answer

What is actually required

The gate exists to answer one question before a workload starts: was this image built by a builder we trust, from a repository we expect, at a commit that passed review? Everything else is packaging.

Measured against that, the policy asserts almost nothing. "A signed attestation exists" says that something inside the organisation emitted a statement about this image. The claim lives in the predicate fields, and the policy never reads them.

The one change that matters

Assert on the predicate, not on its presence:

  • builder.id must be in an allow-list of hosted builders.
  • the source repository in the attestation must be the one that owns this workload.
  • the commit must be reachable from a protected branch at build time.

Without the builder allow-list, a self-hosted runner satisfies the policy, and a self-hosted runner is a machine a team controls. The attestation then reads "built by a machine belonging to the person who wants this deployed", which is SLSA Build L1 wearing an L3 shirt: L1 only requires that provenance exists and permits it unsigned, L2 requires a hosted platform to generate and sign it, and L3 requires a hardened platform where a build cannot forge its own provenance. Self-hosted runners are not disqualified by nature; they are disqualified when anyone can add one and the verifier cannot tell the difference.

What I would remove

  • The signing key in the cluster secret store. Any workload or operator able to read that secret can mint attestations that the gate will accept, so the trust root sits inside the blast radius it is supposed to protect. The cluster should hold only public verification material and the policy that names acceptable identities, with the private key never leaving the build platform.
  • The 94% pass rate as evidence. It measures conformance to a format. The number that would mean something is rejections on predicate fields, and zero in two quarters is consistent with both "nobody tried" and "the gate cannot reject".

What I would leave alone

  • Evaluation at admission. A control that refuses the request beats a report that someone reads later, and this is the one part of the design that is unambiguously right.
  • One blocking rule rather than a library of advisory ones. A single rule everyone understands and nobody can wave through is a stronger control than 240 rules mostly set to warn.
  • Attestation retention, for as long as the release is supported: the audit question arrives after the pipeline that produced the answer has been rebuilt.

How I would argue this in the review

Not "we are insecure". Name the attack the control currently stops, and if nobody can state one, the control is costing pipeline time to buy an audit sentence. Then offer the cheap proof: build an image on a laptop, attest it with a key read out of the cluster secret, and try to deploy it. A negative test takes an afternoon, and running it quarterly matters because the day someone adds a runner pool is the day the gate quietly changes meaning.

When not to tighten it this way

Predicate assertions cost something. The gate now depends on the verification path, so a fail-closed policy turns an outage of the attestation service into an organisation-wide deployment freeze while an incident runs in production, and a stale builder identity in the allow-list blocks a legitimate release at 03:00.

With one team, one registry and one repository, restricting who may push to the registry achieves more, with no new trust root and no new failure mode. Where you do keep the gate, set the failure mode per policy class: block unverified images on the serving path, let a cluster add-on through with an alert, because the alternative is a cluster you cannot repair.