practice

Public Incident Narrative

also called Public Postmortem

An externally published account of an outage that transfers the diagnostic search rather than a conclusion - the timeline that was read, the hypotheses that were wrong and the practice gap that allowed it.

redditincident-communicationpostmortemwritingtrust

A status page reads: "A configuration issue affected some users. The issue has been resolved and we apologise for the disruption." Nobody can act on that sentence, including the company's own engineers eighteen months later, when a similar failure arrives and the only record is a thirty-word apology.

A public incident narrative is the opposite artefact. It states the mechanism precisely enough that a stranger can check their own system against it this week, and it names the practice gap that allowed the failure rather than only the trigger that fired it. The conclusion is cheap to publish and nearly useless; the search is expensive and is the part that transfers.

Why it matters

Outages are the cheapest research the industry has, and they are mostly wasted. A narrative converts one organisation's bad day into a checkable question for everyone else: is our restore path exercised at production scale? beats any amount of generic advice about backups.

It also improves the internal artefact. A write-up intended for publication forces the timeline to be rebuilt from logs rather than from memory, which routinely surfaces that detection lagged customer impact by many minutes and that two recorded hypotheses were wrong. Those are the facts the action list should come from.

Implementation patterns

  • A timestamped timeline recording what was believed at each point, not only what was true. The gap is where the detection work lives.
  • Name the mechanism exactly — setting, version, selector — because nobody can check a paraphrase.
  • State the practice gap separately from the trigger. "An upgrade removed a label" is the trigger. "The configuration was hand-applied, in no repository, by people who have left" is the gap, and the gap is the part other organisations share.
  • Separate the fix list, with owners and dates, so the document is both a story and a commitment.
  • One owner for the artefact, a target of two to four weeks, and a redaction rule agreed with security and legal in advance rather than per paragraph.

Industry example

On 14 March 2023 Reddit was down for 314 minutes after an upgrade from Kubernetes 1.23 to 1.24 stopped applying the old control-plane node label that Calico's route-reflector selector matched, which broke route distribution across the cluster. Reddit published the account on its engineering blog as "You Broke Reddit: The Pi-Day Outage". Two admissions did the heavy lifting: the route-reflector configuration had been applied by hand years earlier through a vendor tool, existed in no repository, and its authors had left; and the control-plane restore they fell back on, there being no supported downgrade, had been exercised only on small clusters rather than at production scale. Both are checkable by any reader against their own estate, which a root-cause statement about a label rename would not have been.

Failure scenarios

  • The apology with no mechanism, which costs the same calendar time and transfers nothing.
  • The sanitised version implying one careless operator, which damages the culture it was meant to protect and guarantees the next write-up is defensive.
  • Publishing a live exposure — a control that can still be bypassed — because the narrative was reviewed for accuracy and not for risk.
  • One deep narrative then silence across the next three incidents, which spends trust rather than building it.

Trade-offs

Choose Gains Pays
Deep public narrative credible trust · industry reuse · better internal artefact days of senior time · legal review · a standing expectation of that depth
Short public statement cheap · low risk no trust gained · nothing learned outside
Internal-only depth full candour · no redaction cost no external credit · less rigour without an outside reader

When not to use it

When the mechanism implicates a customer or a partner, since you cannot name their gap and a narrative with that hole reads as evasion. When the detail is an unclosed exposure, where the shape can be published and the coordinates cannot. When a regulator's filing governs, in which case the engineering narrative stays internal. And when the organisation cannot sustain the depth: pick a level you can repeat for every incident of that severity.

Interview question

Q: You have lost 40 minutes of customer-submitted data in an incident caused by an untested restore path. What do you publish, and what do you deliberately leave out?

What a strong answer covers: publish the timeline with beliefs, the exact mechanism, the fact that the restore path was untested at production scale, the scope of the data loss in plain terms, and the fix list with dates. Leave out individual names, any still-exploitable detail, and speculation about a vendor's internals. A strong answer commits to an update cadence rather than a resolution time while the incident is live, and separates the regulatory notification, which has its own deadline and audience, from the engineering narrative.

Quick check

Quiz: Which part of an incident write-up transfers value to other organisations? — The practice gap and the diagnostic search, not the trigger; unowned configuration and an unexercised restore are everywhere.

Flashcard: What made Reddit's 2023 Pi-Day write-up useful rather than merely transparent? — It named unowned hand-applied configuration and a restore path never run at production scale, both checkable by any reader.