practice

Failure Thinking

also called Pre-mortem, Failure Analysis

The habit of asking what happens when each part fails, applied systematically during design rather than during the incident.

meta-skillsresiliencerisk

The distinguishing habit of experienced architects, and the cheapest to adopt because it requires no tooling — only the discipline of asking the question at every dependency.

The systematic version walks each element and asks what happens when it is down, and separately when it is slow, which is the harder and more common case. Slow is worse than down: down fails fast and callers move on, while slow holds resources, exhausts pools and cascades. A team that has tested for unavailability and not for degradation has tested the easy case.

Then the questions that follow from the answer. Who notices, and how? What does the user see? What recovers automatically and what needs a human? If the answer to recovery is a runbook, has anyone run it?

The pre-mortem is the group version and is unusually effective: assume the system has failed badly six months after launch, and ask everyone to write down why. Framing it as a certainty rather than a possibility gives people permission to voice concerns they would otherwise soften, and it consistently surfaces risks that a risk-register exercise does not.

The related habit for existing systems: read other organisations' public incident write-ups and ask whether the same mechanism exists in yours. The failure modes recur across the industry, and it is considerably cheaper to learn them from someone else's outage.