concept

Build Cache Poisoning

also called Cache Entry Substitution, Poisoned Action Cache

An anti-pattern in which a shared build cache accepts writes from more parties than may change the code - so arbitrary outputs can be substituted under a key and link into everyone's builds with no source change.

build-cacheremote-cachesupply-chainanti-patternslsa

A remote build cache maps an action key — the hash of the command line, declared inputs, toolchain and environment — to the outputs of running that action. A build that finds the key skips the work and takes the stored bytes.

That makes write access a far stronger permission than it looks. Whoever can write a value for a key that others read can place arbitrary bytes into their builds, with no pull request, no change to the repository and no code-owner review. The compromise travels through an artifact store rather than through source control, which is where nobody is looking.

Why it matters

A compromised third-party build step affects one pipeline. A poisoned cache entry affects every build that shares the key, which in a monorepo is every service depending on that target. The blast radius is the dependency graph.

It also defeats the controls teams believe cover this. Source review, protected branches, signed commits and two-person review all operate on the repository, and nothing in the repository changed. SLSA's Build L3 names the missing property: a hardened platform where one build cannot influence another.

Implementation patterns

  • Writes only from trusted builders; everyone else reads. Developer machines get a read-only credential, and this single change removes the class.
  • Writer identity recorded per entry, so a bad entry is attributable to one account rather than to 400 of them.
  • A double-build verifier: rebuild a sample of actions from scratch with the cache disabled and compare outputs byte for byte. This is the only control that detects a substitution after the fact.
  • Capacity sized for a cold farm, because every build goes cold at once when the cache is unreachable and a farm sized for 90-second builds has roughly a tenth of the workers that 14-minute builds need.

Industry example

Treat this as an archetype rather than a named incident: an IDE vendor with a large shared build farm, where several hundred engineers and the CI fleet share one cache because the alternative was a 14-minute inner loop. Everyone evaluates the read path; the write path carries the risk.

The documented analogue is the class of compromised-build-step incidents since 2021, where a downloaded script ran in production pipelines with a job's full credentials. A cache is the same mechanism with wider fan-out: the untrusted contribution is bytes rather than code, consumed by every later build instead of one.

Failure scenarios

  • The unreproducible binary. A service ships an artifact nobody can rebuild from a clean checkout, found weeks later by someone bisecting an unrelated bug.
  • Integrity theatre. The build verifies the downloaded artifact against the hash on its entry and passes, because the writer chose both the bytes and the hash.
  • Signature laundering. The release artifact is signed and the gate verifies it, so a poisoned output acquires a valid provenance chain.

Trade-offs

Choose Gains Pays
Read-only for developers the injection channel closes local builds no longer warm the cache so first local builds stay slow
Writer identity plus sampled rebuilds detection and attribution a verification farm and work to make sampled actions deterministic
No shared cache at all no shared trust boundary the full build cost on every machine and every pipeline run

The honest position for most organisations: CI writes, developers read, and a sampled double build runs weekly. The cost is a slower first local build and a modest verification bill.

When not to use it

The anti-pattern label applies to the write policy, not to caching. Keep a shared cache when builds are long and the team is large. Skip the shared cache entirely when a clean build is already under about 2 minutes, because the infrastructure, its outage behaviour and its trust boundary cost more than they save, and a local cache plus a warm base image captures most of the benefit. Open writes are defensible in one case: a cache scoped to one person or one short-lived branch, where the only reader is the writer.

Interview question

Q: Your remote cache cut median build time from 14 minutes to 90 seconds and every engineer has read-write access. Convince me that is a problem, then tell me what you would change this week and what you would do only if an incident forced it.

What a strong answer covers: the action-key mechanism and why a write is an injection; that source-side controls do not apply because the repository is unchanged; that content addressing proves integrity only relative to a key the writer chose; the one-week change of revoking developer write access and sizing the farm for a cold burst; and the condition that flips the decision, a build short enough that the cache is not worth operating.

Quick check

Quiz: Why does verifying a downloaded cache artifact against the hash recorded on its entry fail to stop poisoning? Answer: the writer chose the bytes and the hash together, so the check proves integrity relative to the key and says nothing about the author's honesty.

Flashcard: Which permission does a shared build cache quietly grant? — The ability to inject bytes into other people's builds without changing any source, so cache write access must match code write access.