On 14 March 2023 Reddit was down for 314 minutes after an upgrade from Kubernetes 1.23 to 1.24 broke how Calico selected its route reflectors. Reddit published a detailed public write-up rather than a short root-cause statement. What did that choice buy them, and where would copying it be a mistake?
Show the full answer Hide the answer
The situation they were in
The trigger was a label. Kubernetes 1.24 stopped applying the old control-plane node label, and Calico's route-reflector configuration still selected on it, so the cluster's route distribution lost its reflectors and pod networking came apart. The configuration had been applied by hand years earlier through a vendor tool, lived in no repository, and the engineers who set it up had left. With no supported downgrade path, recovery meant restoring the control plane from backup — a procedure that had been exercised on small clusters and not at production scale. Reddit came back after 314 minutes, and published the account under the title "You Broke Reddit: The Pi-Day Outage" on its engineering blog.
What the write-up did that a root-cause statement cannot
A root-cause statement transfers a conclusion. A narrative with the timeline transfers the search — which signals were read, which hypotheses were wrong and how long each cost. That is the part other organisations can act on, because the label rename is specific to one upgrade while configuration with no owner and no repository is in nearly every estate.
Two admissions carry most of the value. The first is that a restore path exercised only on small clusters is an untested restore path, which is checkable by any reader this week. The second is that the configuration's authors had left, which reframes the incident from an upgrade mistake into a knowledge-decay problem with a named remedy: find the hand-edited settings, put them in version control, and attach them to the upgrade checklist.
What it cost them
Writing it. A narrative of this kind is days of senior engineering time plus legal and communications review, and it publishes the exact operational detail an attacker enjoys reading. It also sets an expectation: having written once at this depth, a thin statement next time reads as concealment. The cost is a recurring obligation, not a one-off.
Where copying it would be a mistake
- When the mechanism implicates a customer or a partner. You cannot name the third party's gap, and a narrative with that hole in it is worse than a summary.
- When the detail is a live exposure — a bypassable control or an unpatched fleet. Publish the shape and withhold the coordinates until it is closed.
- When the organisation cannot sustain it. A four-engineer team that publishes one deep narrative and then goes quiet for the next three incidents has spent trust rather than built it. Pick a depth you can repeat.
- When the audience is regulated. In some sectors the public artefact must be the regulatory filing and the engineering narrative stays internal, however much the team wants to share it.
Common weak answers
- "Transparency builds trust." True and not the lesson. Trust came from a checkable mechanism, not from tone. An apologetic post with no label name and no restore detail buys nothing.
- "They should have tested the upgrade in staging." Staging had the same hand-edited configuration or none of it; either way the label selector only breaks where the reflectors actually run. The structural fix is configuration in version control with the upgrade reading it, not more staging.