advanced 4 min answer Multiple choice

A fleet of 180000 building controllers all authenticate to the platform with one shared enrolment secret compiled into the firmware. That secret has just appeared in a public teardown writeup. The devices are on cellular links, about 6% are offline at any moment, and a failed credential change means a site visit costing more than the controller. Which migration sequence do you run?

device-provisioningcredential-rotationsecure-elementmigrationblast-radius
Pick one
Show the full answer Hide the answer

The deciding property

The compromised secret is also the only thing that currently authenticates the channel you need in order to fix it. Every sequence that removes it before a replacement exists on the device locks out the devices it was supposed to protect, and the lockout is a truck roll per unit.

So the sequence must create the new credential through the compromised channel, while narrowing what that channel is allowed to do. Narrow first, replace second, revoke third.

The sequence

  1. Narrow the secret's authority today, before any firmware ships. It should authorise exactly one operation: enrol one device identity, once, at a rate the platform caps per serial-number block. No telemetry, no commands, no configuration reads. This is a server-side change, takes hours, and shrinks the blast radius of a secret that is already public.
  2. Add detection on the narrowed path. Enrolments from unexpected serial ranges, a second enrolment for a serial already enrolled, and geographic or volume anomalies. You now see exploitation rather than inferring it later.
  3. Ship firmware that generates a keypair inside the secure element - private key non-exportable - emits a certificate signing request, authenticates that request with the old secret plus whatever device-unique material exists, and stores the issued certificate. Keep the old secret working for this device until the new certificate has been used successfully at least once.
  4. Switch each device to certificate authentication and record the switch in the registry as the authoritative fact, not as a log line.
  5. Revoke the shared secret per cohort, once that cohort's enrolment rate has flattened. Cohort by region or firmware version so that a mistake costs one cohort.
  6. Time-box an amnesty for the long tail, then produce a site-visit list. The list is the deliverable, not the failure.

Where it can diverge and how you would know

Double enrolment is the main hazard: a device that enrols, loses the response, and enrols again leaves two valid certificates for one physical unit. Make enrolment idempotent on a device-supplied request identifier, and alert on any serial with more than one unrevoked certificate. The registry must be able to answer "how many distinct credentials can speak as this device" for every unit, because that is the number an auditor and an attacker both care about.

The point of no return and how long it really takes

Revocation per cohort is the irreversible step; everything before it is reversible by leaving the old path open. With 6% offline at any moment, enrolment follows a decay curve: most devices in days, 95% in two to three weeks, and a residue that includes units that are powered down seasonally. Plan on a month of dual-credential operation and a permanent, audited break-glass enrolment path for devices that return after the amnesty, used with manual approval and alerting.

Why the other options fail

  • Revoke now and require certificates on next connection. The device has no code that can produce a certificate, so its next connection fails, and the firmware that would fix it is delivered over the channel you just closed. This is the classic chicken-and-egg lockout, and it converts a credential incident into a 180,000-site field operation.
  • Rotate the secret and re-sign the firmware. You pay a full firmware cycle and still hold one secret that unlocks the fleet and is present in every unit sold. The next teardown writeup resets the clock.
  • Push certificates from the manufacturing database. A private key that can be pushed is a private key that exists off the device, which makes the manufacturing database a fleet-wide skeleton key and the push channel a key-distribution channel. The device must generate its own key and never surrender it.
  • SIM allow-list and keep the secret. A SIM identifier is an identifier, not a credential; it is presented by the network rather than proven by the device, and it survives being swapped into other hardware. This buys a filter and leaves the authentication broken.

When not to run this sequence

If the hardware has no secure element or keystore, stop and re-plan, because a per-device key held in flash is copyable and the migration buys bookkeeping rather than security. In that case the honest options are network-level isolation plus a narrowed enrolment path, or accepting that fleet identity is weak and moving the controls that matter to the account level.

Skip the staged version entirely in the opposite case: a fleet of a few hundred devices with a technician already visiting on a maintenance cycle. Scheduled on-site re-keying is cheaper than a dual-credential window, a cohort plan and a month of monitoring, and it ends with no shared secret in circulation at all.