In 2023 a team added a build rule forbidding the domain module from importing the persistence module, and it went green. In 2026 the domain module imports persistence in fourteen places and the build is still green. No one disabled the rule and no one edited it. Reconstruct how a passing check ended up protecting nothing.
Show the full answer Hide the answer
The trigger
A fitness function is, in Ford, Parsons and Kua's 2017 formulation, a mechanism that gives an objective assessment of an architectural characteristic. The assessment is only as good as the set of things it looked at, and a rule that looked at nothing reports success. Four mechanisms produce that outcome, and in a three-year-old codebase more than one is usually present:
- Scope by string. The rule matches
com.acme.domain... Someone renames the package tocom.acme.core.domainduring a refactor, the rule's matcher selects zero classes, and the assertion over an empty set passes. - A frozen baseline. Tools that support recording existing violations start with 6 and end with 14, because "add to the baseline" is one line in a pull request and removing the import is a day of work.
- A stage nobody watches. The rule was moved out of the pull-request gate because it was slow, into a nightly job whose notifications go to a channel the original team no longer reads.
- A new module. The domain was split into
domainanddomain-model; the rule names the first.
Why detection lagged for three years
Every signal available was positive. The build was green, the test count went up, and the architecture document still said the rule existed. There is no metric anywhere that goes red when a check stops checking, because the output of a check is a boolean about the code, not a statement about itself.
The precedent is older and much more expensive. Knight Capital's 1 August 2012 deployment went to seven of eight order-routing servers; the eighth still carried code retired in 2003 behind a flag the new code reused, and the firm lost more than $460 million in roughly 45 minutes. A safeguard's presence is not its operation, and nothing in either system was capable of reporting the difference.
The tempting local fix, and the structural one
The local fix is to repair the matcher and clear the baseline. Necessary, and it buys about eighteen months before the next rename.
The structural fix is to make every architecture rule assert its own coverage. Three moves, in order:
- Fail on an empty selection. The rule declares a minimum class count and fails below it, so a rename breaks the build instead of silencing it.
- Keep a known violator. A permanently non-compliant fixture class that the rule must flag, verified in the same run. If the rule ever stops flagging the fixture, the rule is broken. This is the only construction that tests the test.
- Put an expiry on every baselined violation. A recorded exception with a date fails the build when the date passes. Without expiry, a baseline is a way to convert a rule into a comment.
Then move the rule back into the pull-request gate. A check whose result nobody sees within the hour is documentation.
The general lesson
Automated governance introduces a failure mode that manual review does not have: manual review fails loudly by not happening, and automation fails silently by happening vacuously. The trade is worth making, and it is only worth making if the check is instrumented as a system component rather than trusted as a fact.
When not to bother with this rigour
For a rule that was advisory in the first place — naming conventions, package layout preferences — coverage assertions cost more than the rule is worth, and a green build that means nothing is an acceptable outcome. Reserve the meta-checking for rules whose violation is expensive to reverse: dependency direction, data residency boundaries, anything that decides where customer records may be written.