Attestation Quarantine Path
also called Restricted Admission Path, Attestation Fail-Safe Lane
A restricted network path that a device failing attestation is admitted to - permitting only enrolment, time sync and firmware download - so that an integrity check cannot lock a fleet out of the channel used to repair it.
A fleet requires a signed boot measurement on every connection and admits only measurements on an allow-list. Then a bad release ships, devices revert to the previous image exactly as designed, and that image's digest was pruned from the list last month. The recovery mechanism is now the thing keeping the fleet offline, and it fails on precisely the units that already proved they cannot run the new build.
The quarantine path is the answer to this class of lockout. Attestation failure does not close the door; it narrows what is behind the door. A device that cannot prove its integrity is admitted to a lane that offers enrolment, time synchronisation, firmware download and a diagnostic report endpoint - and nothing else. No telemetry is accepted, no commands are issued, no customer data is served.
Why it matters
Integrity checks are the only controls whose failure mode can be worse than the attack they prevent. A unit that cannot connect cannot be updated, logged into, measured or recovered, and for a device behind a locked plant room that means a truck roll costing more than the hardware. The expected cost of a strict door is the probability of a policy mistake times the fleet size times the site-visit cost, and policy mistakes are not rare: a pruned digest, a silicon revision that changes a measurement for new stock, a clock that resets and makes certificate validation impossible.
Quarantine converts an irreversible outcome into a reversible one, which is the property that makes strict attestation deployable at all.
Implementation patterns
- Append-only allow-list with expiry per digest rather than deletion, and the recovery and factory images' digests are never expired. Deletion is the single most common cause of self-inflicted lockout.
- Publish the measurement before the image. Enforce the ordering in the build pipeline, so a promoted build whose digest is unpublished cannot reach devices. A runbook step here fails on the day it matters.
- A separate, narrow authorisation scope for the quarantine lane, issued as a token valid for 15 minutes rather than days, so a quarantined device cannot sit in the lane indefinitely harvesting firmware.
- Attestation gates authorisation scope, not connectivity. This is the design rule the pattern encodes: the measurement decides what the session may do, which is adjustable, rather than whether a session exists, which is not.
- Rate limit and alert on the lane. Quarantine volume is a security signal: more than roughly 0.5% of the fleet in quarantine means either a bad release or an attacker probing with modified firmware, and both need a page within minutes.
- Break the clock deadlock explicitly. Allow an unauthenticated, signed time response in the lane, since a device with a reset real-time clock cannot validate your certificate and therefore cannot reach anything else.
- Signed diagnostic builds, so a field engineer has a path that does not require defeating secure boot.
Industry example
Android's Verified Boot is the platform example closest to the pattern: a device whose boot state fails verification boots into a restricted or warned state that still allows recovery, rather than refusing to boot at all, and the state is reported to anything that asks. The confidential-inference designs published in 2024 show the strict end of the same spectrum - a client refuses to send data to a server node that cannot attest to a specific software image, and checks that measurement against a public append-only transparency log first. The append-only property is the relevant part here: measurements accumulate rather than being replaced, so an older approved image stays verifiable and nothing can quietly drop a build from the record.
Failure scenarios
- Pruned rollback digest. The previous image cannot attest, so automatic revert produces an offline device instead of a working one.
- New silicon, same software. A boot-ROM or fuse revision changes the measurement for units built after some date; new stock fails to enrol while the fleet dashboard looks healthy, and the cause is a supplier change nobody told the platform team about.
- A quarantine lane that is too generous. It accepts telemetry or exposes a configuration API, so an attacker running modified firmware gets a usable channel and the attestation requirement bought nothing.
- A quarantine lane with no alert. Devices accumulate in it for weeks; the fleet is in a degraded state that no dashboard describes.
- Clock deadlock. The lane requires TLS validation that the device cannot perform, so devices with reset clocks are locked out of the mechanism designed to rescue them.
Trade-offs
Quarantine costs a second admission path with its own authorisation model, its own tests and its own abuse surface - and the honest admission that a device you cannot vouch for still gets to talk to you. Every additional capability in the lane erodes the guarantee, so the lane must be reviewed with the same care as the main path.
What it buys is the ability to run strict attestation without betting the fleet on the allow-list being right forever. Without it, the maximum safe strictness is whatever your release process can guarantee, which in practice is lower.
When not to use it
If devices are physically controlled and the realistic threat is network-side, secure boot plus mutual TLS is sufficient and a quarantine lane is machinery without a job. Skip it too where a failed device is trivially recoverable by hand - a rack in your own data centre, a developer kit on a desk - because the lockout cost that justifies the lane is a site visit you cannot afford. And if the deployment genuinely cannot tolerate an unverified device on the network at all, do not build a soft lane: build a physically separate recovery process and fund the visits, rather than shipping a lane whose scope will widen under pressure.
Interview question
Q: You require fresh attestation on every connection. Walk me through what happens the first time you need an emergency firmware rollback, and what you would have built beforehand.
What a strong answer covers: the pruned-digest lockout and why it hits the devices least able to recover; an append-only allow-list with expiry and permanently retained recovery digests; publish-measurement-before-image as a pipeline invariant; the quarantine lane's exact scope and its short-lived token; the clock deadlock and the signed time response that breaks it; quarantine volume as a paging signal; and the framing that attestation should gate authorisation scope rather than connectivity, because scope is reversible under pressure and connectivity is not.
Quick check
Quiz: Why must an attestation allow-list be append-only with expiry rather than edited in place? Because a removed digest makes the image that carries it unable to connect - and the most likely such image is the rollback or factory build that recovery depends on.
Flashcard: What does a device that fails attestation get to do? Enrol, sync time, download firmware and report diagnostics, under a short-lived narrow token - and nothing else - so an integrity failure is a degraded session rather than a site visit.