metric

Glossary Term Half-Life

also called Definition Decay Rate, Term Attestation Age

The time after which half a glossary's definitions no longer match the SQL that implements them, measured by sampling rather than assumed, and used to set the review budget a glossary can actually afford.

business glossarydefinitionsdecayattestationreview cadence

A glossary is written in a quarter of enthusiasm and read for years. Nobody plans for the gap. The definition of "active customer" is agreed in March, the model that computes it acquires a new filter in July when a trial tier is launched, and the glossary entry stays exactly as it was, because changing SQL and changing an entry in a governance tool are two different jobs owned by two different people.

Half-life makes that gap measurable. Sample 20 governed terms at random, open the model each one claims to describe, and count how many still match on grain, filter and time basis. Do it twice, six months apart, and you have a decay curve. Most estates that have never measured find their first sample somewhere between 55% and 75% accurate, which puts the half-life in the region of one to two years, and it is shorter in domains where the product changes fast. Check the entry against what is in production, not against what the owner believes is in production.

The number is not interesting on its own. It is interesting because it converts an unbounded ambition, "keep the glossary current", into a budget: a glossary can carry only as many terms as its half-life and its review capacity jointly allow.

Why it matters

A stale glossary is worse than no glossary, and the mechanism is specific. An entry with a governance badge is read as a warranty by people who cannot inspect the SQL: analysts new to the company, auditors, and anyone building a downstream metric. They do not verify it, because the badge exists so that they do not have to. So a wrong entry propagates into work rather than being caught, and when it is finally discovered the damage is not one report but everything built on it.

The half-life also settles an argument that governance programmes have every year: how many terms to govern. If re-attesting a term costs 30 to 45 minutes and the half-life is 18 months, the review cadence must be well inside 18 months, which for 150 terms is roughly 90 hours a year of protected time. Ten times that many terms is not ten times the effort; it is a programme that quietly stops reviewing and keeps publishing.

Implementation patterns

  • Sample, do not survey. Twenty randomly chosen terms checked properly beats asking 200 owners whether their entries are current, because self-report measures confidence rather than accuracy.
  • Record an attestation date on every entry and show the age on the page. A reader can then discount a two-year-old definition without anyone declaring it wrong.
  • Set the cadence from the measured curve, not from a governance calendar. A domain with a nine-month half-life gets reviewed twice a year; a stable regulatory domain with a five-year half-life gets reviewed when something changes.
  • Trigger on lineage, not only on the clock. When the model implementing a term changes in a way that touches the columns the definition names, drop the entry to review-pending and send the owner the diff. Event-driven review is cheaper because the reviewer arrives holding the change.
  • Publish the half-life itself next to the glossary. It is the honest statement of what the artefact is worth, and it makes the case for the review budget without an argument.

Industry example

The pattern is visible in reverse in every organisation that has run a definition reconciliation after an audit finding. A characteristic archetype: a subscription business with 1,400 harvested candidate terms and 260 declared governed. A 20-term sample found 11 still matching their models. Three of the mismatches were the same change, a trial tier added in a product launch that altered what counted as active, and none of the three owners had been told the launch touched their definition. The programme's response was not to write more entries but to cut the governed set to 90 and wire the remainder to lineage events.

Failure scenarios

  • The glossary is measured once, reported as 68% accurate, and never measured again. A single point is not a half-life and cannot set a cadence.
  • The sample is drawn from the terms the owners nominate, which are the ones they just reviewed. Random selection is the whole method.
  • Attestation becomes a button. Owners click "still correct" without opening the model, and accuracy falls while the dashboard turns green. The counter-measure is to sample attested entries too, and to publish the rate at which attested entries fail the sample.
  • The half-life is used to justify abandoning the glossary. The decay is a property of the business changing, not of the practice failing.

Trade-offs

Choose Gains Pays
Measure the half-life A defensible governed-term budget and a cadence with a reason A day of careful sampling twice a year and an uncomfortable first number
Assume terms stay current No measurement cost A warranty on entries nobody has checked since they were written

The uncomfortable part is political rather than technical. The first measurement tells a sponsor that the artefact they funded is substantially wrong, which is why programmes avoid measuring it and why the measurement is the thing that makes the rest of the practice honest.

When not to use it

Below roughly 40 governed terms with a single data team, skip the metric. The owners can check the whole set in an afternoon, so sampling estimates a number you could count exactly. It also does not apply to terms fixed by external definition, a regulatory reporting field whose meaning is set by a supervisor's taxonomy does not decay with your product, and forcing it into the same cadence wastes the review budget that the volatile terms need.

Interview question

Q: Your company's glossary has 300 governed terms and the sponsor believes it is current. You suspect it is not. How would you find out, what would you do with the answer, and how would you stop the same decay recurring?

What a strong answer covers: random sampling against the implementing models rather than surveying owners; checking grain, filter and time basis rather than reading the prose; two measurements to get a rate rather than a point; converting the rate into a review budget and cutting the governed set to fit it; and moving from calendar review to lineage-triggered re-attestation so the reviewer sees the specific change.

Quick check

Quiz: Why does sampling 20 random terms beat asking all 300 owners whether their entries are current? Self-report measures the owner's confidence, not the entry's accuracy, and the owners most likely to respond are the ones who just reviewed theirs.

Flashcard: A glossary reports 68% of definitions still matching their models. What can and cannot you conclude? You have one point, not a rate. A second measurement six months later gives the half-life, which is what sets the review cadence and caps how many terms the programme can carry.