Reddit published a postmortem for its 14 March 2023 outage: a Kubernetes upgrade from 1.23 to 1.24 on its oldest and largest cluster renamed a node label, the CNI's route reflectors selected nodes by the old label and found none, pod networking collapsed, and the site was down for 314 minutes before a restore from backup. Review the *postmortem* as a document. What makes it useful, and what would you push back on in any version of it?
Show the full answer Hide the answer
What is actually required of a postmortem
Not a narrative. A postmortem has to change future behaviour, which means a reader who was not there can answer three questions: what the system did, why that was possible, and what is different now. Everything else is detail competing with those answers.
What makes this one useful
- The causal chain is specific enough to generalise from. "A renamed label emptied a selector" is a mechanism, and a reader can go and check their own selectors this afternoon. A postmortem's value is the actions it provokes in other teams, and vague ones provoke none.
- It names the latent condition, not just the trigger. The trigger was an upgrade. The latent condition was an implicit dependency on a label's name, undeclared and unvalidated, in the oldest and least-understood cluster. That distinction is the whole craft: the trigger was ordinary and will recur, so the finding has to be about the condition.
- It is honest about duration and about the recovery choice. 314 minutes with the cause found after the restore is an uncomfortable fact, and publishing it is what makes the rest credible.
- It explains why diagnosis was hard, which is the part that actually determines incident duration and the part most postmortems omit entirely.
What I would push back on in any postmortem of this shape
- "We will be more careful with upgrades" as an action. Unfalsifiable and therefore not an action. The acceptable form names a mechanism: a pre-upgrade check that asserts every selector still matches a non-empty set, run automatically.
- Action items without owners, dates and a verification step. An action nobody owns is a sentence. Each should say who, by when, and how we will know it worked — ideally "this check now fails the upgrade", which is testable.
- No statement of what would have detected it sooner. Time-to-detect and time-to-diagnose belong in the timeline as explicit numbers with their own remediations. If the honest answer is "nothing we had would have caught this", that is the single most important finding in the document.
- Blast radius treated as given. The report says this was the oldest, largest cluster carrying the legacy core, while newer service layers stayed up. That is the structural finding: a dependency graph with one node whose loss takes most user flows. The remediation is not about upgrades at all, and an upgrade-focused action list quietly accepts the topology.
- No "what nearly saved us". Near-misses and the thing that was almost tried are where the cheapest improvements hide.
- Counting the restore as success. Restoring from backup without knowing the cause is a legitimate call under pressure and leaves you unable to prevent a recurrence. The document should say whether the cause would have been found without the restore, and what would make the next diagnosis faster.
What I would leave, even though it looks odd
The length and the technical depth. The instinct in a review is to ask for an executive summary and fewer internals. Resist it for an incident of this size: the detail is the reusable part, and the audience that changes its behaviour as a result is engineers who need the mechanism. Add a summary at the top; do not shorten the body.
Also leave the absence of names. A blameless account reads as evasive to people unfamiliar with the practice. It is correct — the engineer who ran the upgrade is not the finding — and defend it on the ground that naming individuals suppresses the reporting the next postmortem depends on.
When not to write one at all
Not every incident earns a full postmortem, and pretending otherwise is how the practice decays. A document per minor degradation produces a backlog nobody reads and actions nobody does, which teaches the organisation that postmortems are paperwork. Set a threshold — customer-visible impact beyond some duration, data loss, or a repeat of a previous cause — and for everything below it, a few lines in the incident ticket. Spend the review capacity on the incidents that will teach something, and the repeat-cause clause is what stops small recurring problems slipping under the bar forever.