Review this decision record. It chooses a single Postgres primary over sharding for an orders service measured in production at 1200 writes per second, lists three rejected alternatives with reasons, and ends with "Revisit in 12 months". What would you change, what would you leave alone, and what one line makes the decision reviewable?
Show the full answer Hide the answer
What is actually required
A decision record has two jobs. The first is to explain the choice to someone reading it later, and this one does that: a measured baseline, three alternatives, the reasons for rejection. The second job is to let a future reader decide whether the decision is still correct without re-running the analysis, and this record cannot do it. "Revisit in 12 months" fails for two separate reasons, and both matter.
It has no owner, so the calendar entry belongs to nobody and the review does not happen. More importantly, a date is not evidence. The decision can become wrong in month three if a large customer is onboarded, and can still be right in month twenty if growth stalls. A date-triggered review arrives at a time uncorrelated with the thing that would make the decision wrong.
What I would leave alone
The three rejected alternatives with their reasons. Reviewers often want these cut as noise; they are the only artefact that stops the organisation re-proposing a rejected option every six months. I would also leave the 1,200 writes per second alone and resist the urge to pad it with a forecast. A measurement with a date beats a projection, and the projection is what the trigger is for.
The one change that matters
Replace the revisit date with a falsification trigger: the observation that makes the decision wrong, with a threshold, a signal and an owner.
This decision is wrong if sustained write throughput on
ordersexceeds 4,000 writes/s, or p99 commit latency exceeds 25 ms while throughput is under 3,000 writes/s, or the working set exceeds 60% of instance memory. Signal: the existingorders-dbdashboard. Alert at 70% of each threshold. Owner: the orders tech lead.
Three properties make that line work. It fires on the quantity that actually drove the choice, which forces the author to name it — and the common discovery is that they cannot, which tells you the decision rested on taste. It fires early enough to act, because the alert is at 70% of the threshold, and sharding an orders service is on the order of two quarters of work; a trigger that fires at the point of pain is a trigger that fires too late. It is checkable from telemetry alone, so the reviewer does not have to read the prose to know the answer.
What it costs
Roughly an extra half hour per decision, and a harder conversation: naming a threshold means accepting that you can be shown to be wrong by a number, in public, on a dashboard. That is the real reason records say "revisit in 12 months". The second is maintenance: a trigger wired to a dashboard that gets renamed is dead, so the quarterly job is checking that each live signal still exists. Ten decisions is about an hour a quarter.
When this is the wrong answer
Decisions that are cheap to reverse do not need triggers; the cost of writing one exceeds the cost of simply changing your mind later. Triggers earn their keep on the one-way doors: data models, partition keys, public API shapes, anything with a migration cost measured in quarters. And where the driver genuinely is calendar-shaped — a contract that expires, a vendor end-of-life date, a regulation taking effect — a date is the correct trigger and the only one.
Common weak answers
- "Add a quarterly architecture review that walks all decision records." This scales as records × quarters and produces a meeting whose output is a list of records nobody has new information about. Triggers invert it: the record raises its hand.
- "Write the assumptions section properly." Assumptions are necessary and insufficient. An assumption says what you believed; a trigger says what observation would prove the belief wrong, and only the second one closes the loop.