A fleet platform decides to require a fresh hardware attestation on every device connection. The gateway issues a nonce, the device returns a boot measurement signed inside its secure element, and only measurements on an allow-list are admitted. The team gets replay resistance and firmware integrity. What have they given up, and when does that bill arrive?
Show the full answer Hide the answer
What is gained
A captured measurement stops being useful, because the nonce binds it to one exchange. A unit running modified firmware cannot open a session at all rather than being detected later. And a build discovered to be compromised can be excluded fleet-wide by deleting one digest from the allow-list, which is the fastest containment primitive a fleet can have: minutes, with no firmware push and no device cooperation.
What is paid
- The allow-list becomes a release-blocking dependency, in the opposite order to intuition. The measurement must be published before the image it describes reaches any device. A build promoted to devices ahead of its digest is a build that cannot connect, and the failure is silent on the server side: the device simply never appears.
- Energy and latency on every reconnect. A nonce round trip plus a signature inside a constrained secure element is on the order of tens to low hundreds of milliseconds and a measurable charge draw. A device on a flapping cellular link that reconnects 50 times a day pays that 50 times, which matters on a battery budget and on a metered plan.
- Reconnection storms get more expensive on the server too. The gateway now performs a nonce issue and a signature verification per connection, so 200,000 devices returning after an outage is a cryptographic thundering herd rather than a socket one. Size the verification path, not just the accept queue.
- You cannot debug what you cannot admit. "Boot a patched image and collect logs" is no longer available without a separately signed diagnostic build, and the field engineer discovers this at the site.
- It composes badly with a wrong clock. A device that boots with no trustworthy time cannot validate the gateway's certificate, so attestation sits behind a dependency loop that attestation itself cannot break.
When the bill arrives
At the first emergency rollback. The team prunes old digests for hygiene, a bad release goes out, devices revert to the previous image as designed, and that image's measurement is no longer on the allow-list. The recovery mechanism is now the thing that locks the fleet out, and it fails on exactly the devices that already proved they cannot run the new build.
The second arrival is quieter: a silicon vendor ships a new boot-ROM or fuse revision, identical software produces a different measurement on units built after some date, and new stock fails to enrol while the dashboard shows a healthy fleet.
How to keep the option to reverse
- The allow-list is append-only, with explicit expiry per digest rather than deletion, and the recovery and factory images' digests are never removed.
- Attestation failure degrades to quarantine, not to a closed door. Admit the device to a restricted path that permits only enrolment, time sync and firmware download. Attestation then gates authorisation scope, which is reversible, rather than connectivity, which is not.
- Publish measurement before image, enforced in the pipeline, so ordering is a build-system invariant rather than a runbook step.
When not to require it per connection
If the devices are in physically controlled space and the threat you actually face is network-side, secure boot plus mutual TLS already covers it, and per-connection attestation buys little for real battery and operational cost. The middle position is usually right: attest at enrolment and whenever the firmware measurement changes, then issue a short-lived token carrying the attestation result, so the hot path verifies a token and the expensive proof happens on a schedule you choose. State the token lifetime, because that number is your actual exposure window.
Require it per connection when a single device compromise reaches other customers' data or physical actuation - the case the 2024 generation of confidential-inference designs was published to address, where a client refuses to send data to a node that cannot attest. Most building or metering fleets are not that case.