A company certifies 400 of its 4000 dashboards with a badge an owner and a review at the moment the badge is granted. Eighteen months later an audit finds the certified set has the highest proportion of reports whose underlying model has changed since certification. Review the design. Which single change matters most?
Show the full answer Hide the answer
What is actually required
The badge exists to answer one question for a reader who cannot inspect the SQL: can I quote this number in a decision? That answer is a claim about a point in time. A certification without an expiry asserts something about the past and is read as a statement about the present, which is why the audit found what it found. Certification also draws attention, so certified dashboards are the ones most likely to be extended, refactored and repointed, which makes the certified set drift faster than the uncertified set rather than slower.
The one change that matters
Bind the badge to the lineage. When any upstream model a certified dashboard depends on changes in a way that touches the columns it reads, the badge drops to "review pending" automatically and the owner is notified with the diff. It returns when the owner re-attests.
Two properties make this work where a calendar does not. It is event-driven, so a dashboard untouched for two years is not re-reviewed for the sake of it. And the reviewer arrives holding the specific change, which turns a 20-minute re-read into a 2-minute judgement. At 400 dashboards with a realistic upstream change rate, this produces on the order of a few re-attestation requests a week rather than a quarterly block of about 130 hours.
Why the other options fail
- Raise the bar. Staleness is a function of time since attestation and of upstream churn, not of how hard the badge was to earn. A smaller certified set drifts at the same rate and now covers less, so more decisions are made on uncertified numbers.
- Move certification to the semantic layer. This is a good change for a different problem, the one where two dashboards disagree because each wrote its own SQL. It does not fix this: a dashboard can consume a governed metric and still apply the wrong filter, the wrong date grain or point at a deprecated model. Certify both, but the layer alone does not expire anything.
- A quarterly review meeting. This is the same control with a calendar attached. The reviewer has no signal about what changed, so the review degrades into re-confirming from memory, and 400 dashboards at 20 minutes each is roughly 133 hours a quarter that nobody has.
- A usage leaderboard. Usage is a retirement signal, not a trust signal. An unused stale certified dashboard harms nobody; a heavily used one is the whole problem, and the leaderboard actively protects it.
When this is the wrong answer
With fewer than about 30 certified assets and one data team, lineage-triggered expiry is infrastructure for a problem you can solve by reading a changelog. Announce model changes in a channel the five owners read. The automation earns its cost when the number of certified assets exceeds what one person can hold in their head, roughly the point where nobody can name the dependants of a model from memory.