intermediate 2 min answer

Forty teams have been writing architecture decision records for six years at roughly three per team per year. Estimate the size of the corpus and the fraction that still describes the system as it stands today. What does the number rule in or out about how the log should be structured?

decision-records-writingestimationcorpus-rotsupersededindexing
Show the full answer Hide the answer

The assumptions, stated

  • 40 teams × 3 records per team per year × 6 years = 720 records, written at a steady 120 a year.
  • Each record describes a decision with a finite life. The rate at which a live decision is overturned — by a re-platform, a vendor change, a team merge, a boundary move — is the number that dominates everything below. Call it the obsolescence hazard h per year.
  • A record written k years ago is still accurate with probability (1 − h)^k.

The arithmetic, shown

Surviving records = 120 × Σ(1 − h)^k for k = 1…6.

  • At h = 15% (a stable platform): 0.85 + 0.72 + 0.61 + 0.52 + 0.44 + 0.38 = 3.63 → about 435 live, roughly 60%.
  • At h = 30% (an organisation that re-platformed once in the period): 0.70 + 0.49 + 0.34 + 0.24 + 0.17 + 0.12 = 2.06 → about 245 live, roughly one third.

So the honest answer is 700-odd records of which somewhere between a third and 60% are still true, and the band is wide because the hazard rate drives it. Everything else — the exact write rate, the team count — moves the answer by tens of records. The hazard rate moves it by two hundred.

What the number rules out

It rules out reading the log. At 700 entries, answering "what is our current position on asynchronous messaging" by browsing means opening on the order of twenty documents and determining, unaided, which of them still matches what is in production. The expected cost of a lookup is tens of minutes, so engineers stop looking and ask in chat instead — which is the observed end state of most decision logs and the reason people conclude the practice fails.

It rules in three cheap structural fixes:

  • A status field on every record — proposed, accepted, superseded-by — maintained at the moment of supersession. That costs about 2 minutes per overturned decision, roughly 40 edits a year at h = 30%, against tens of minutes per reader lookup. The arithmetic is not close, and the field cannot be reconstructed later because the people who knew have moved on.
  • An index per component rather than per date. Decisions are looked up by subject, so a flat chronological folder is the wrong access path from about 150 entries onward.
  • A revisit condition in the record itself, so a reader can tell whether the decision was overtaken without reconstructing history.

The rule: a decision log is a database with no index until someone adds one, and it crosses from browseable to unusable at roughly 150 entries — about four years for a 40-team organisation, which is why the problem always appears to arrive suddenly.

When this is the wrong answer

For a single team with 30 records, all of this is overhead: chronological order is fine and everyone remembers the live decisions. Add the status field from day one anyway, since it is free at write time and cannot be reconstructed later.