metric

Review Capacity Budget

also called Board Throughput, Governance Capacity

The number of design reviews a governance body can actually complete in a period, which determines the significance threshold rather than being determined by it.

architecture-review-boardsgovernancequeueingthresholdsautomation

A board of five meets for 90 minutes a week and is asked to review all significant changes across 120 teams. Nobody does the arithmetic, so the threshold is set by ambition and the queue sets the real policy.

Capacity: 90 minutes × 45 weeks ≈ 4,050 minutes, which at 25 minutes of board time per review is about 160 reviews a year. Demand, at one architecturally significant decision per team per quarter, is 480 a year, or 12,000 minutes. The board can see roughly a third of the decisions it has been asked to govern, on an optimistic estimate of demand.

Why it matters

A queue that receives ten items a week and clears three and a half does not degrade gracefully; it grows without bound, and the wait reaches six weeks within two months. Every familiar pathology follows from that arithmetic rather than from the board's behaviour: teams starting work before approval, approval becoming a formality, and the board acquiring a reputation as an obstacle while believing it is under-resourced.

Stating the capacity first inverts the conversation. The threshold becomes a consequence of arithmetic rather than a matter of taste, and the discussion moves from "what is significant" to "which 160 decisions do we most want a human conversation about".

Implementation patterns

  • Compute capacity before setting the threshold, and publish both. The threshold is then defensible to a team that wants a review and to an executive who wants more coverage.
  • Define significance by reversibility and blast radius, with worked examples: a new primary datastore, a new jurisdiction for personal data, a new third party on the critical path, a change to an interface other teams consume, anything creating a new always-on cost line.
  • Automate the residual. The 320 decisions the board stops seeing become machine-checked policy in the pipeline, which runs on all of them at near-zero marginal cost. This is the only mechanism that scales with 120 teams, and the board's job becomes writing the checks.
  • Make approval advisory with a recorded deviation route. A team may proceed over an objection by recording the decision and its rationale, which removes the board from the critical path and makes its influence rest on argument quality.
  • Measure queue wait, not meetings held. Wait time is the number teams experience and the leading indicator of the board being routed around.
  • Timebox review to the decision, not the design. A 25-minute slot forces a one-page submission, which is also better for the author.

Industry example

The published alternatives to the queue model all reduce board-mediated reviews in favour of recorded decisions and automated checks: lightweight architecture decision records as the default artefact, technology radars as published guidance rather than per-case approval, and policy-as-code gates in deployment pipelines. In the other direction, the UK Government Digital Service's spend-control gate from 2011 is a documented case of a central approval threshold that worked, and it worked because the threshold was explicit, monetary and narrow rather than a general claim on all significant change.

Failure scenarios

  • The six-week queue, with work started before approval, so the review becomes a retrospective justification.
  • Rubber-stamping under load, where the board clears its backlog by approving quickly and the quality signal disappears while the process metric improves.
  • Routing around, where teams reclassify their change as insignificant, which is rational and unobservable.
  • Board growth as the fix, doubling capacity, multiplying coordination cost, and leaving the order of magnitude unchanged.
  • Reviewing designs rather than decisions, which triples the time per item and produces feedback the author cannot act on.

Trade-offs

Narrowing the threshold means genuinely significant decisions will be made without review, and some of them will be wrong. That is the cost, and it should be stated plainly rather than hidden behind an aspiration to review everything. What you buy is a board whose reviews are timely and whose objections carry weight, plus automated coverage of the decisions it no longer sees. The alternative is a board that nominally covers everything and effectively covers nothing on time.

When not to use it

In a regulated environment where the review exists to produce evidence, the cost of the review is the point, and the threshold should be sized to the evidential requirement instead of to engineering risk. In a 15-team organisation the arithmetic is comfortable and the board can genuinely see everything, so the budget is not binding and imposing a threshold removes value. The metric matters at the scale where demand exceeds capacity by an order of magnitude, which is most organisations above about 40 teams.

Interview question

Q: You take over an architecture review board with a six-week queue and a poor reputation. Leadership asks you to reduce the wait without reducing governance. What do you do in the first month, and what do you tell leadership about what they are giving up?

What a strong answer covers: computing and publishing capacity before touching the process · a threshold based on reversibility and blast radius with examples, and the explicit admission that decisions below it will be unreviewed · policy-as-code for the residual, as the honest answer to "without reducing governance" · the shift to advisory approval with a recorded deviation route so the board leaves the critical path · measuring queue wait rather than meetings held · and refusing the apparent fix of enlarging the board, with the arithmetic to justify the refusal.

Quick check

Quiz: A board clears 3.5 reviews a week against 10 arrivals. What happens to the wait, and what is the only fix that changes the order of magnitude? — The queue grows without bound and the wait reaches six weeks within two months; only reducing arrivals through a narrower threshold, with automation covering the rest, changes the magnitude.

Flashcard: Why should capacity be computed before the significance threshold is written? — Because capacity is fixed and demand is a policy choice; computing it turns the threshold into arithmetic rather than a matter of taste.