An IDE vendor with a large shared build farm gives 400 engineers read and write access to the same remote build cache that CI uses, cutting median build time from 14 minutes to 90 seconds. Months later one service starts shipping a binary that nobody can reproduce from a clean checkout. Which change removes the largest class of risk?
Show the full answer Hide the answer
The deciding property
A remote build cache maps an action key — the hash of the command line, the declared inputs, the toolchain and the environment — to the outputs of running that action. A build that finds the key skips the work and takes the stored bytes.
So the deciding question is: who may write a value for a key that everyone else reads? With 400 writers, any one of them can place arbitrary bytes under the key for a compile step in a service they have never worked on, and every later build of that service links them in without executing a single line of attacker-controlled code locally. The cache is a code-injection channel that bypasses source review, code owners and branch protection, because nothing was changed in the repository.
This is the same shape as a compromised third-party build step, with a wider blast radius: a build step affects one pipeline, a poisoned cache entry affects every build that shares the key. SLSA's Build L3 names the property directly — a hardened platform where one build cannot influence another.
Why the other options fail
- Clear the cache nightly. It bounds exposure to a day and leaves the mechanism intact. Re-poisoning costs one build, and the daily clear also throws away the hit rate that justified the cache, so you pay the full cost and keep the risk.
- Verify the artifact against the entry's recorded hash. This is the instinct of anyone who has debugged a corrupted download, and it detects corruption in transit. It cannot detect substitution, because the writer chose both the bytes and the hash. Content addressing proves integrity relative to the key, never the honesty of the key's author.
- Pin the toolchain. Correct and necessary, for a different bug: an incomplete key makes two genuinely different builds collide on one entry, which produces wrong results with nobody lying. Fixing key completeness stops accidents and does nothing about a writer who intends to lie.
- Sign the release artifact. The signature binds a publisher to bytes that were already poisoned upstream, so the gate happily verifies a signed compromised binary. Signing answers "who published this", not "what went into it".
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| A verifier double-builds a sample of actions from scratch and compares outputs | developer writes for a low-risk subset | an independent rebuild detects mismatch within hours rather than never |
| Builds are hermetic and the cache is content-addressed per action with signed entries per writer identity | writer identity recorded per entry | the cache becomes attributable and a bad entry is traceable to one account |
| The project's clean build is already under two minutes | drop the shared cache | the infrastructure and its failure modes cost more than they save |
When not to run a shared cache at all
A remote cache is a production dependency of every build. When it is slow or unreachable, every build is cold at once: the farm that was sized for 90-second builds now needs the capacity for 14-minute builds, which is the same thundering herd a cache-flush causes in any serving system. If your build is minutes rather than tens of minutes, local caching plus a warm base image buys most of the benefit without adding a shared trust boundary.