Evidence ledger
One row per claim in The renewal failed a month before anyone noticed: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Topic: how production systems manage the lifetime of the credentials machines use to authenticate to each other, why expiry still causes outages at organisations that know better, and what the forced move to 47-day certificates changes.
All links fetched 2026-08-29. One row per claim. Quotes are copied, not paraphrased.
| # | Org | Title | Tier | Published | Checked | URL | Claim I take from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|
| 1 | Bazel (Google) | Postmortem for *.bazel.build SSL certificate expiry | postmortem | 2026-01-16 | 2026-08-29 | https://blog.bazel.build/2026/01/16/ssl-cert-expiry.html | Automation existed, failed for a month, and told nobody | auto-renewal began failing around 2025-11-26 because a removed subdomain was still on the certificate; "These failures never triggered any notifications" |
| 2 | Bazel (Google) | same | postmortem | 2026-01-16 | 2026-08-29 | https://blog.bazel.build/2026/01/16/ssl-cert-expiry.html | The detection gap, not the expiry, set the outage length | expired 2025-12-26 07:35:39 UTC; users reported at 09:00; a responder started investigating at 15:08; new certificate deployed 20:31; about 13 hours |
| 3 | Bazel (Google) | same | postmortem | 2026-01-16 | 2026-08-29 | https://blog.bazel.build/2026/01/16/ssl-cert-expiry.html | A DNS change two months earlier was the real trigger | docs-staging.bazel.build DNS records removed 2025-10-30; the Google-managed certificate still listed it, and every domain must be reachable for renewal to succeed |
| 4 | DigiCert | Bug 1910322: random value in CNAME without underscore prefix | postmortem | 2024-07 to 2024-08 | 2026-08-29 | https://bugzilla.mozilla.org/show_bug.cgi?id=1910322 | A validation defect can force a fleet-wide rotation on 24 hours' notice | 83,267 certificates, 6,807 subscribers, 172,047 unique domains; "The 24-hour revocation rule applies in this circumstance as the issue impacts domain validation." |
| 5 | DigiCert | same | postmortem | 2024-07 | 2026-08-29 | https://bugzilla.mozilla.org/show_bug.cgi?id=1910322 | Subscribers who could not rotate in time went to court | a customer obtained a Temporary Restraining Order on 2024-07-30 blocking revocation; DigiCert: "We did receive a TRO in connection with this revocation" |
| 6 | DigiCert | same | postmortem | 2024-07 | 2026-08-29 | https://bugzilla.mozilla.org/show_bug.cgi?id=1910322 | The CA named inability to rotate as an infrastructure risk | "many other customers operating critical infrastructure, vital telecommunications networks, cloud services, and healthcare industries are not in a position to be revoked without critical service interruptions" |
| 7 | DigiCert | same | postmortem | 2024-07 | 2026-08-29 | https://bugzilla.mozilla.org/show_bug.cgi?id=1910322 | The defect was introduced by a 2019 refactor and found by an outsider five years later | "One path through the updated system did not automatically add the underscore nor check to see if the random value had a pre-appended underscore"; reported by a third party on 2024-07-03 |
| 8 | Let's Encrypt | Incident 2026-05-08, status page | postmortem | 2026-05-08 to 2026-05-13 | 2026-08-29 | https://letsencrypt.status.io/pages/incident/55957a99e800baa4470002da/69fe2d6698ca07050eb4b1b3 | A CA can stop issuing, and a fleet on short lifetimes is then on a clock | 18:37 UTC "We have been made aware of a potential incident and are shutting down all issuance."; resumed 21:03 UTC |
| 9 | Let's Encrypt | same | postmortem | 2026-05-13 | 2026-08-29 | https://letsencrypt.status.io/pages/incident/55957a99e800baa4470002da/69fe2d6698ca07050eb4b1b3 | ARI is the mechanism that makes a mass reissue survivable | "We have identified that our cross-certified subordinate CAs are missing Extended Key Usage (EKU) fields which are now required. We are revoking and reissuing our cross-signs"; "Our ACME Renewal Information API is signalling affected certificates to renew now" |
| 10 | Microsoft Azure | PIR, tracking ID ML7_-DWG | postmortem | 2025-12 | 2026-08-29 | https://azure.status.microsoft/status/history/?trackingId=ML7_-DWG | Rotation itself is a failure mode, not only a control | an "inadvertent automated key rotation" caused ARM authorization failures; Cosmos DB accounts were "unintentionally configured with the option to automatically rotate the keys enabled" while ARM used keys in manual mode |
| 11 | Microsoft Azure | same | postmortem | 2025-12 | 2026-08-29 | https://azure.status.microsoft/status/history/?trackingId=ML7_-DWG | Two clouds failed three hours apart and the first did not save the second | keys for Azure Government and Azure in China were created the same day in February 2025, "starting completely independent timers"; the team "did not have a sufficient understanding of the inadvertent key rotation to be able to prevent impact in the second sovereign cloud" |
| 12 | CA/Browser Forum | Ballot SC-081v3: schedule of reducing validity and data reuse periods | adr | 2025-04-11 | 2026-08-29 | https://cabforum.org/2025/04/11/ballot-sc081v3-introduce-schedule-of-reducing-validity-and-data-reuse-periods/ | Public certificate lifetime falls to 47 days on a fixed schedule | reduction "from 398 days to 47 days" starting March 2026 and concluding March 2029; SAN validation data reuse "from 398 days to 10 days" |
| 13 | CA/Browser Forum | same | adr | 2025-04-11 | 2026-08-29 | https://cabforum.org/2025/04/11/ballot-sc081v3-introduce-schedule-of-reducing-validity-and-data-reuse-periods/ | The stated rationale is staleness, and the vote was unopposed | "Certificates are representations of a point in time state of reality... The more time passes from that moment of issuance, the more likely it becomes that data represented in the certificate diverge from reality."; issuers 25 yes / 0 no / 5 abstain, consumers 4 yes; proposed by Clint Wilson (Apple) |
| 14 | Google Chrome Root Program | Moving Forward, Together | adr | current, retrieved 2026-08-29 | 2026-08-29 | https://chromium.googlesource.com/website/+/5943ad5e004c1cde2f869f1d67008a729e9ec12e/site/Home/chromium-security/root-ca-policy/moving-forward-together/index.md | The lifetime cut is explicitly a forcing function for automation | "Reducing certificate lifetime encourages automation and the adoption of practices that will drive the ecosystem away from baroque, time-consuming, and error-prone issuance processes." |
| 15 | Google Chrome Root Program | same | adr | current | 2026-08-29 | https://chromium.googlesource.com/website/+/5943ad5e004c1cde2f869f1d67008a729e9ec12e/site/Home/chromium-security/root-ca-policy/moving-forward-together/index.md | ACME is becoming a root-store entry requirement | "over 50% of the certificates issued by the Web PKI rely on ACME"; intent that "all Chrome Root Store applicants must be part of PKI hierarchies that offer ACME services" |
| 16 | IETF / ISRG | RFC 9773, ACME Renewal Information (ARI) extension | adr | 2025-06 | 2026-08-29 | https://www.rfc-editor.org/rfc/rfc9773.html | The CA, not the client, should decide when renewal happens | "Allowing issuing CAs to suggest a period in which clients should renew their certificates enables dynamic time-based load balancing."; clients "Select a uniform random time within the suggested window" |
| 17 | IETF / ISRG | same | adr | 2025-06 | 2026-08-29 | https://www.rfc-editor.org/rfc/rfc9773.html | ARI exists partly to make mass revocation survivable | "a CA could suggest that clients renew prior to a mass-revocation event to mitigate the impact of the revocation" |
| 18 | IETF | RFC 8739, Short-Term, Automatically Renewed (STAR) certificates in ACME | adr | 2020-03 | 2026-08-29 | https://datatracker.ietf.org/doc/html/rfc8739 | The formal slack rule: publish the next credential by half-life | "the next certificate MUST be made available by the ACME CA at the URL... halfway through the lifetime of the currently active certificate at the latest" |
| 19 | IETF | same | adr | 2020-03 | 2026-08-29 | https://datatracker.ietf.org/doc/html/rfc8739 | Short lifetimes trade revocation risk for availability risk, and the RFC says so | "expiration replaces revocation"; "Reducing the caching properties... makes STAR clients increasingly dependent on the ACME server availability"; effective lifetime "no longer than 4 days" for web use |
| 20 | Stanford / CMU | Topalovic, Saeta, Huang, Jackson, Boneh. Towards Short-Lived Certificates | paper | 2012 (W2SP) | 2026-08-29 | https://www.ieee-security.org/TC/W2SP/2012/papers/w2sp12-final9.pdf | The origin of the whole trajectory, and the slack rule stated in 2012 | "The Online Certificate Status Protocol (OCSP) is as good as dead."; "We suggest a certificate lifetime as short as four days"; "If this fetch fails, the web site is not harmed since the certificate obtained the previous day is active for a few more days giving the administrator and the CA ample time to fix the problem." |
| 21 | ISRG | Aas et al. Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web (CCS '19) | paper | 2019-11 | 2026-08-29 | https://www.abetterinternet.org/documents/letsencryptCCS2019.pdf | 90 days was chosen as a forcing function, not a security optimum | "Let's Encrypt limits certificate lifetimes to 90 days... This is short enough to strongly encourage operators to automate things (few people will want to manually renew certificates every quarter) but long enough to allow for manual renewal if necessary" |
| 22 | ISRG | same | paper | 2019-11 | 2026-08-29 | https://www.abetterinternet.org/documents/letsencryptCCS2019.pdf | Automation measurably removed the lapse problem | "Prior to Let's Encrypt, renewal lapse occurred for about 20% of trusted certificates"; with LE, "few Let's Encrypt renewals (2.9%) occur after the prior certificate has expired"; 2.2% of Top Million LE sites serve expired certificates against 3.9% of all HTTPS sites |
| 23 | ISRG | same | paper | 2019-11 | 2026-08-29 | https://www.abetterinternet.org/documents/letsencryptCCS2019.pdf | Scale of the automated issuance model | "issued over 538 million certificates for 223 million domains" in just over three years |
| 24 | Let's Encrypt | Ending Support for Expiration Notification Emails | blog | 2025-01-22 | 2026-08-29 | https://letsencrypt.org/2025/01/22/ending-expiration-emails | The human safety net was deliberately removed | Josh Aas: "Over the past 10 years more and more of our subscribers have been able to put reliable automation into place for certificate renewal."; service ended 2025-06-04; cost "tens of thousands of dollars per year" |
| 25 | Let's Encrypt | We Issued Our First Six Day Cert | blog | 2025-02-20 | 2026-08-29 | https://letsencrypt.org/2025/02/20/first-short-lived-cert-issued | The operational rule at six days: renew three times per lifetime, check ARI daily | Josh Aas: "short-lived certificates should be renewed every two to three days"; "ARI checks should happen at least once per day"; rationale "certificate revocation doesn't work very well" |
| 26 | Uber | Our Journey Adopting SPIFFE/SPIRE at Scale | blog | 2023-11-09 | 2026-08-29 | https://www.uber.com/en-IE/blog/our-journey-adopting-spiffe-spire/ | Internal machine identity at scale is an issuance-capacity problem before it is a security problem | "4,500 services running on hundreds of thousands of hosts across four clouds"; each host caches "thousands of SVIDs"; an LRU cache let them register "around 2.5 times more workloads" and cut "CPU usage by 40%" |
| 27 | Netflix | Introducing Lemur | blog | 2015-09-21 | 2026-08-29 | https://netflixtechblog.com/introducing-lemur-ceae8830f621 | The ownership problem was named a decade ago and is unchanged | Glisson, Chan, Hagen: "Certificates have expiration dates — if they are allowed to expire without replacing communication can be interrupted, impacting a system's availability."; "There are quite a few steps to this process and much of it is typically handled by humans." |
| 28 | Keyfactor | Key takeaways from the 2024 PKI and Digital Trust Report | blog | 2024-04-08 | 2026-08-29 | https://www.keyfactor.com/blog/key-takeaways-from-the-2024-pki-digital-trust-report/ | Detection, not renewal, is where the hours go | "On average, it takes 2.6 hours to identify a certificate-related outage and another 2.7 hours to remediate it."; "It takes an average of eight staff members to remediate an outage."; "The average organization experienced nine certificate-related incidents over the past 12 months."; "Only 32% of organizations reported using a dedicated certificate lifecycle management tool." |
| 29 | cert-manager | FAQ: renewal timing and terminal failure | source | current | 2026-08-29 | https://cert-manager.io/docs/faq/ | The default slack is one third of the lifetime, and failure is a field nobody watches | "If renewBefore has not been set, Certificate will be renewed ⅔ through its actual duration."; terminal failure is visible only as "Issuing condition set to false and has status.lastFailureTime set"; retries back off "1 to 32 hours by default" |
| 30 | cert-manager | Issue #6378: renewal fails during issuer downtime, continues to fail after recovery | source | 2023-09-28 | 2026-08-29 | https://github.com/cert-manager/cert-manager/issues/6378 | Renewal automation can latch into a failed state that outlives the fault | backoff widens 1 hour to 4 hours to 13+ hours; no new CertificateRequest is generated; manual deletion is required to recover |
| 31 | cert-manager | PR #7399: add renew window to restrict when certificate renewal can happen | source | 2024-11-05, closed unmerged 2026-05-29 | 2026-08-29 | https://github.com/cert-manager/cert-manager/pull/7399 | Operators want change-freeze windows for renewal, and it took 18 months to land | release note: "Add field renewTimeWindow that let's the user restrict when the certificate can be renewed with a cron expression"; motivation was renewal "outside business hours"; closed with "This feature is now implemented in #8258" |
| 32 | cert-manager | PR #8258: certificate renewal policy and windows | source | 2025-11-16, merged 2026-03-23 | 2026-08-29 | https://github.com/cert-manager/cert-manager/pull/8258 | The redesign that did land | "This PR arises from the requirement to allow disabling or scheduling of certificate renewals. We use cron to allow mentioning the start of window and duration" |
| 33 | Kubernetes SIG Cluster Lifecycle | cluster-api issue #10522: cert-manager rotation may lead to webhook downtime up to 90s | source | 2024-05 | 2026-08-29 | https://github.com/kubernetes-sigs/cluster-api/issues/10522 | Rotation is a change, and changes have a window | webhook failures observed for about 49 seconds; "it can take up to 60-90s for the new secret data to be propagated to the pod, while requests already try to get validated with the new secret"; "usually happens every 60 days" |
| 34 | ISRG / IETF | ACME Renewal Information, IETF 121 slides (Aaron Gable) | talk | 2024-11-06 | 2026-08-29 | https://datatracker.ietf.org/meeting/121/materials/slides-121-acme-ari-and-profiles-00 | The working-group record for ARI and certificate profiles | slide deck for draft-ietf-acme-ari-06, presented by Aaron Gable of Let's Encrypt; covers Retry-After advice to servers and the ACME profiles proposal |
| 35 | TechCrunch | Here's what caused yesterday's O2 and SoftBank outages | blog | 2018-12-07 | 2026-08-29 | https://techcrunch.com/2018/12/07/heres-what-caused-yesterdays-o2-and-softbank-outages | The largest published blast radius for a single expiry, reported rather than first-party | about 32 million UK subscribers affected; Ericsson CEO Börje Ekholm: "The faulty software that has caused these issues is being decommissioned and we apologize not only to our customers but also to their customers." |
Tier mix
postmortem 11, adr 8, paper 4, blog 6, source 5, talk 1. 35 rows over 28 distinct artefacts, 18 distinct hosts.
Where the record runs out
- No first-party Ericsson account survives online. The December 2018 outage is the largest published blast radius in this class and the only accounts still reachable are journalism quoting an Ericsson statement. Treat the mechanism as reported, not documented.
- The talk layer is genuinely thin. There is no equivalent of the postmortem corpus in conference form. The design argument for this topic happens in CA/Browser Forum ballots, IETF drafts and bug trackers, not on stage. If you want the reasoning, read the ballots.
- Nobody has published a measured before-and-after for internal PKI automation. Uber and Netflix describe what they built; neither publishes the expiry-incident rate before and after, so there is no public figure for what the investment buys.
- The 47-day transition has no operational evidence yet. The 200-day tier began on 2026-03-15, five months before this guide. There is no published account of an organisation running a large estate through the 100-day or 47-day tiers, because nobody has.
- The survey numbers are vendor-collected. Keyfactor sells certificate lifecycle management. The figures are the best public ones available and they are not independent.