advanced 3 min answer

A travel marketplace in the mould of Expedia reports on its central architecture board: 212 submissions last year, 97% approved, a 3-day median to decision and a reviewer satisfaction score of 4.2 out of 5. Leadership asks whether the board changes anything. Which numbers would you pull first, in what order, and which of the four reported figures is misleading?

architecture-reviewsmetricsgovernancegoodhartcoverage
Show the full answer Hide the answer

The first three things I would look at, and why in that order

  1. Material change rate. Sample 20 approved submissions, read the design as submitted and the system as built, and count the ones where the board caused a change a reader can point at — a datastore swapped, a boundary moved, a rollback path added. This is the only direct measure of effect.
  2. Coverage. Count the significant changes the organisation actually made from production evidence rather than from submissions: new datastores created, new external egress destinations, new public endpoints, new runtimes introduced. Compare that to 212. A board reviewing a third of the real change set cannot be a gate no matter how good the reviews are.
  3. Submission timing. The fraction of submissions that arrived before implementation started. A design submitted after the code exists cannot be changed by a review, so those rows inflate every other number while contributing nothing.

The diagnosis

Approval rate is not a gate metric, and 97% is consistent with both a board that is excellent and a board that is ornamental. A healthy review mostly approves, because teams that know the bar meet it before submitting. The distinguishing number is whether the approved design differs from the submitted one.

Planning rule from practice: under roughly 10% material change, the board costs more than it returns and should become advisory unless something else justifies it. 20–40% is a working review. Above 60%, the bar is being discovered in the room rather than published, and the fix is review-readiness criteria, not more reviewers. The failure this board is exposed to is not a bad design slipping through a review; it is a bad design never reaching one.

The misleading signal

Satisfaction of 4.2 is the trap. Teams rate a fast approval highly, so satisfaction correlates with throughput and not with value. The 3-day median is the same mistake one layer down: fast is good when the review has substance and is exactly what you would see if the board were a stamp. Both numbers measure the experience of the queue, not the output of it.

The fix

Publish the short list of decisions that genuinely need central sign-off — anything creating a new datastore, a new destination for personal data, a new public endpoint or a new production language. Everything else goes to a peer review with a recorded decision and a named approver. Then report two numbers only: coverage of that short list, and material change rate. Cut submissions by roughly half and the per-submission effect should rise, which is the point.

The alert worth having is on coverage: a new datastore or a new egress destination appearing in production without a matching submission, detected from infrastructure state, not from self-reporting. Choose that alert over any dashboard of review throughput.

When this is the wrong answer

In a newly-assembled organisation after an acquisition, a near-100% approval rate with low change rate can be the correct state for a year: the board's job is to build the shared picture, and the output is a map and a vocabulary rather than design changes. Measure it then on whether reviewers can describe each other's systems, and switch to change rate once the map exists.