Governance Metric Capture
also called Proxy Target Capture, Single-Metric Governance
What happens when a proxy measure becomes a published target - teams move the measure by the cheapest available path, which is rarely the path that improves the property it stood for.
An engineering organisation of 26 teams publishes one governance metric: change failure rate. Within a quarter it has improved by 40%, and deployment frequency has fallen by half.
Nothing dishonest happened. The cheapest way to reduce the share of changes that fail is to make fewer, larger changes, so teams batched work, and each remaining deploy carried more risk, took longer to verify and was harder to roll back. The property the metric stood for — the ability to change the system safely — got worse while the number got better, and the organisation celebrated.
Governance metric capture is the general case. A proxy is chosen because it is measurable and correlates with something that is not. Publishing it as a target changes behaviour, and the behaviour it changes is whatever moves the proxy most cheaply, which is almost never the mechanism the designer had in mind. The correlation that justified the proxy is the first thing to break.
Why it matters
Architecture governance runs on proxies: incident count, change failure rate, mean time to restore, test coverage, percentage of services on the golden path, number of open findings. Each is reasonable as an observation and each has a cheap path that does not involve improving anything.
- Incident count falls when the definition of an incident narrows.
- Mean time to restore falls when the clock starts later, at declaration rather than at first impact.
- Test coverage rises with tests that execute code and assert nothing.
- Golden-path adoption rises when teams copy the template and diverge immediately, so adoption climbs while the share of services running in production more than two template versions behind climbs with it.
- Open findings fall when findings are closed as "accepted risk".
The damage is not the gamed number; it is the loss of the signal. After capture, the metric no longer tells anyone whether the property is improving, so the organisation has paid for a measurement system and is now flying blind while believing it is instrumented.
Implementation patterns
- Never publish a target without a paired counter-metric, and read them only as a pair. Change failure rate with deployment frequency. Incident count with time-to-detect. Cost per request with p99. Coverage with escaped-defect rate.
- Target the pair, not either number. "Reduce change failure rate without reducing deploy frequency" is harder to game because the cheap paths move the two in opposite directions.
- Watch for divergence as the alarm. If one of a pair moves more than about 20% while the other is flat or moving the wrong way, suspect the measurement before believing the improvement.
- Separate diagnostic metrics from targeted ones, explicitly and in writing. The same number is safe on a dashboard and dangerous in a quarterly objective, and teams need to know which it is.
- Version the definition and announce changes. Most capture happens through quiet redefinition, so a definition with an owner and a change log removes the easiest path.
- Spot-check the underlying property twice a year by a different method — read 10 incident reports, review 10 "covered" tests — because the only reliable detection is an independent look.
Industry example
The four delivery metrics popularised by the DORA research programme, reported in its annual State of DevOps publications through the 2010s and 2020s, are deliberately paired for this reason: throughput measures — deployment frequency and lead time — sit alongside stability measures — change failure rate and time to restore — and the reported finding is that strong performers improve both together. Teams that adopt one half as a target reliably move it by damaging the other, which is the most common misuse of the set in practice. The same shape appears wherever a platform group is measured on golden-path adoption: adoption rises, and the share of services more than two template versions behind rises with it, because copying the template counts as adoption and staying current does not.
Failure scenarios
- Definition drift, where the metric improves because the counting rule changed and nobody logged it.
- Reclassification, the cheapest form: a severity-2 incident recorded as a degradation.
- Batching, which improves any per-change metric by reducing the number of changes.
- Local optimisation at a cost elsewhere: a team hits its latency target by shedding the slow 1% of requests, which lands as a support problem in another department.
- The metric outliving its reason, still reported and targeted years after the property it proxied stopped being the constraint.
Trade-offs
Pairing metrics costs clarity. One number per quarter is easy to communicate and a pair invites argument about which matters more, which is exactly the argument the pair exists to force. Publishing nothing avoids capture entirely and also removes the ability to see whether anything is improving, which is worse.
Choose a single headline metric only where the cheap path to moving it is the path you want, and that condition is rare enough to be worth checking explicitly before publishing anything.
When not to use it
This is a concern about targets, not about measurement. Instrument freely; a number on a dashboard that nobody is rewarded for does not get captured, and withholding measurement to avoid capture is the wrong lesson.
It also matters less where the metric is a direct measure rather than a proxy. Revenue, error budget consumed and actual cloud spend are the things themselves, and while they can still be moved by undesirable means, the proxy failure mode does not apply. The anti-pattern is specific to a measurable stand-in for an unmeasurable property, and that is where the counter-metric is mandatory.
Interview question
Q: Your CTO wants one number on a slide each quarter to show that engineering is getting safer, and has proposed change failure rate. Argue for it, then tell them what you would insist on alongside it and what you would do in the first quarter it improves by 30%.
What a strong answer covers: why the metric is reasonable and cheap to collect · the cheap path that moves it — fewer, larger changes — and why that makes the system less safe · the paired counter-metric, deployment frequency, with the target expressed over the pair · divergence as the alarm rather than the level · a documented definition with an owner so improvement by redefinition is visible · and the independent check in the quarter it improves, namely reading a sample of change records to confirm the batch size did not grow.
Quick check
Quiz: You publish change failure rate as the single governance metric for 26 teams. What improves and what gets worse? The metric improves because teams batch into fewer larger deploys; deployment frequency falls and the ability to change safely — the property the metric stood for — gets worse.
Flashcard: What single rule prevents governance metric capture? — Never publish a target without its paired counter-metric and target the pair, because the cheap paths to moving one push the other the wrong way.