Treat an operator's attention late in an incident as a hard constraint rather than a behaviour to coach. In January 2017 a GitLab engineer working past the end of their shift ran a destructive command against the primary database host instead of the replica. Which controls change that outcome and which are theatre?
Show the full answer Hide the answer
The trigger, stated precisely
Late on 31 January 2017, around 23:00 UTC, a GitLab engineer troubleshooting replication after their shift had ended removed a PostgreSQL data directory on the primary rather than the replica. Its published postmortem describes roughly 300 GB gone in seconds and about six hours of data lost for good, because the restore that worked came from an LVM snapshot that existed by accident for unrelated staging work.
The control in place at the moment of the action was this: a human distinguishing two visually identical terminals by reading a hostname in a shell prompt, late at night, after their shift.
Why this is a constraint, not a behaviour
Reading-carefully has an error rate, and it is not a constant: it rises with hours on task, with time of day and with the pressure of a live incident, which is exactly when destructive commands get typed. Any control whose reliability is a function of operator attention is least reliable precisely when it is load-bearing.
Constraint thinking treats that like any other capacity limit: you do not fix a disk at 100% IOPS by asking it to try harder, and the binding constraint at 23:00 is attention rather than knowledge. Investment in training, runbook wording and reminders buys nothing against a constraint not made of knowledge.
Controls that change the outcome
- Make the destructive command carry its target. A wrapper that takes the expected hostname and refuses on mismatch turns a misread prompt into a rejected command. Cost: one script. Its zero-cost companion is a red prompt for production terminals.
- Make the irreversible action reversible. Removing a data directory becomes a move to a timestamped trash path on the same volume, reclaimed after 24 hours, so the data is still there during the minute when the mistake becomes obvious. Cost: headroom for roughly one copy of the largest directory, so hold the volume near 50% free. This is the highest-value control, because it converts the error class from permanent to annoying.
- Require a second operator's approval token for irreversible actions on a primary, enforced by the tool rather than the runbook. Cost: real delay during incidents, so scope it to a named list of operations.
Which are theatre
A training reminder, a checklist line reading "verify the host", a policy forbidding late-night operational work, and a second approver on the change document rather than on the command. Each adds a step that depends on the same depleted attention, and each lets the organisation record that it responded. The test: pick your three most destructive commands and ask whether one tired person can execute each in under ten seconds with no second artefact. If yes, the control is the person.
What it costs, and the honest limit
Wrapping destructive operations adds friction for the people best at them, and trash-then-reclaim needs headroom and a reaper that actually runs — an unreaped trash directory filling a data volume is a new outage. Rehearse the reclaim path or you have traded one failure for another.
When the simpler alternative wins
For a single-operator system a coloured prompt and a trash wrapper are the whole programme; approval tooling for a three-person team spends weeks on what an undo already covers. Approval mechanics are justified by the number of hands with production access multiplied by the blast radius of the worst command, not by the severity of the last incident.
Common weak answers
- "The real problem was the broken backups." Both failed, at different layers: backups decide how much you lose, controls on the destructive path decide whether you lose anything. Fixing only the restore path leaves the next tired engineer one misread prompt from exercising it.
- "Nobody should run manual commands on production." Correct as a direction and useless during an unplanned incident, which is when these commands exist. Design for the operator who is already in the shell, because that is the operator you will have.