advanced 3 min answer

The CA/Browser Forum's SC-081v3 ballot cuts the maximum public TLS certificate lifetime to 200 days from 15 March 2026, 100 days from 15 March 2027 and 47 days from 15 March 2029, with domain validation reuse falling to 10 days. Your estate has roughly 900 public certificates, perhaps 40% renewed by hand. Sequence the response.

certificatesacmeautomationestate-wide-changedeadline-driven
Show the full answer Hide the answer

The sequence

The deadlines are external and non-negotiable, which makes this unusually easy to plan and unusually unforgiving to defer. The work is not renewal, it is removing humans from renewal, and the three phases give you a schedule rather than a choice.

  1. Inventory from the network, not from the spreadsheet. Discover certificates by scanning your own public endpoints and by reading certificate transparency logs for your domains. The spreadsheet will be missing the ones that matter: a load balancer configured in 2021, a partner-facing endpoint, a certificate pinned inside a mobile app.
  2. Classify by renewal mechanism, not by expiry date. Three buckets: fully automated (ACME with a tested renewal), automatable (standard endpoints where ACME can be introduced), and structurally manual — hardware appliances, certificates pinned in shipped mobile clients, a partner who requires a certificate file by email, embedded devices that cannot fetch a new chain. The third bucket is the whole problem and it is always smaller than feared and harder than expected.
  3. Automate the automatable immediately, because 200-day certificates already mean a renewal cycle most teams can survive by hand and 47-day ones do not. At 47 days, 900 certificates is roughly 7,000 renewals a year, around 20 a day: manual renewal is arithmetically finished, not merely inconvenient.
  4. Renew at a third of lifetime, not at expiry. With a 47-day certificate, renew at day 16, which leaves two further attempts before an outage. Alerting at seven days to expiry is a habit from the 398-day era and leaves no room.
  5. Attack the manual bucket with architecture, not process. Terminate TLS at a layer you control so the appliance behind it keeps a long-lived private certificate. Replace pinning with a pinned CA or a pin set with backup keys, or remove pinning where the threat model does not justify it. For the partner who needs a file, negotiate an endpoint or accept a standing manual process with two named owners.
  6. Make issuance a platform capability with a monitored queue, so a new service gets a certificate by default and an expiring one raises a ticket against a team rather than against nobody.

Where it can diverge, and how you would know

The divergence that causes the outage is a certificate that renews correctly and is not reloaded by the process serving it. The file on disk is new, the running server still presents the old chain, and nothing notices until expiry. Monitor the certificate the endpoint actually serves, from outside, not the file on disk — this is the single most valuable check in the whole programme.

The second divergence is domain validation. With reuse falling to 10 days, every renewal effectively revalidates, so a DNS configuration that worked when validation was cached quarterly must now work continuously. A broken delegation or a stale CAA record becomes a renewal failure rather than a once-a-year annoyance.

The point of no return

There is none in your control, which is what makes this different from most migrations: the dates arrive whether or not you are ready. Each phase is reversible only in the sense that you can issue shorter-lived certificates early, which is the recommended rehearsal. Pilot 47-day certificates on a real service in 2026, before the phase that compels it, so the failures are found at a time of your choosing.

How long it really takes

For 900 certificates with 40% manual, assume two to three quarters of engineering spread across the teams that own the endpoints, with the structurally manual bucket taking longer than everything else combined. The schedule is set by the slowest owner, which makes this an enterprise architecture problem rather than a platform-team task: the standard, the tooling and the deadline are central, and the remediation is distributed.

When not to centralise this

Do not try to take ownership of every endpoint's configuration into one team. Publish the standard, provide the issuance platform and the external monitoring, and let teams integrate, because a central team holding 900 configurations becomes the bottleneck the compressed schedule cannot afford. Where internal, private PKI is in use, these ballot deadlines do not apply at all, and extending them to internal certificates on principle is work you have invented.