Reddit's published account of its 14 March 2023 outage describes 314 minutes of downtime, says it took the team hours to decide that a rollback was the right course of action, and notes the subtle cause was not found until hours after the site had already been restored from a backup. The restore was the fix. Why did deciding to use it cost more than doing it, and which decision taken before the incident would have changed that?
Show the full answer Hide the answer
The trigger
An upgrade of a large Kubernetes cluster from 1.23 to 1.24 removed a node role label. The cluster's Calico route reflectors were selected by that label, found no nodes, and pod networking collapsed across the cluster. The published write-up gives the outage as 314 minutes and the resolution as a restore from backup.
The important fact for this question is the ordering: service came back first, and the cause was identified hours later. Diagnosis was never on the critical path to recovery. Hours were nonetheless spent on it, and hours more on deciding whether to restore.
Why the decision cost more than the action
Under a deadline, responders are comparing two options with asymmetric information. Continuing to diagnose has a known marginal cost — another fifteen minutes — and an unknown payoff. Restoring has a known payoff and an unknown total cost: how long it takes on this cluster, what state it loses, whether anyone has ever done it here. An option whose price nobody can state cannot be compared to anything, so the group keeps buying the next fifteen minutes instead of choosing, and does so repeatedly, which is how a 45-minute decision becomes a three-hour one.
This is not a failure of nerve. It is the predictable behaviour of a group asked to choose between a small known cost and an unbounded unknown one, and it recurs in every organisation whose recovery path has never been executed.
What makes a recovery option cheap to choose
- A rehearsed, timed number. "We restore this cluster in 55 minutes" turns the unknown into a comparison. A runbook that has never been executed does not price the option; it only describes it.
- A written data-loss bound. "We lose up to five minutes of cluster state" is the second half of the price.
- Pre-delegated authority. The incident commander may call it without finding an executive.
- A stopping rule agreed in advance. "Full outage with no confirmed cause within 45 minutes: execute the restore path." With a 55-minute restore, the worst case becomes a 100-minute outage, which is a different incident from a 314-minute one.
- Evidence capture as step zero of the restore. Snapshot the broken state and copy logs off before rebuilding, so choosing the fast path does not cost the postmortem.
The structural fix versus the tempting local fix
The tempting fixes are "test upgrades better" and "write a runbook". Both are reasonable and neither touches what happened here, because the delay was a decision, not a defect and not a missing document. The structural fix is to price and rehearse the recovery option so it can be chosen in one minute, and to pre-commit the rule that triggers it. The signal that it worked is not fewer incidents; it is a shorter gap between the start of an outage and the first irreversible recovery action.
When this is the wrong answer
Restore-early is wrong where the fault is silent corruption rather than a hard outage. Restoring an unknown corruption re-creates it, and the state you destroy is the only evidence of what was wrong. The stopping rule therefore needs two clauses: hard outage, restore early; suspected corruption or data divergence, freeze the state and diagnose, accepting a longer degradation to keep the evidence. Deciding which clause applies is itself a judgement that should be made once, in writing, not at 02:00.