Decision Latency
also called Time to Architectural Decision, Decision Queue Time
The elapsed time from a design question being raised to a decision being made, whose length changes not only when designs ship but how complicated they are - because optionality is the rational hedge against slow decisions.
An architecture forum meets weekly with six twenty-minute slots. About fourteen cross-team interface requests arrive each week, so the backlog grows by eight a week and within a month a new request waits roughly four weeks.
The delay is the obvious cost. The expensive cost is what teams do about it. They bundle unrelated changes into one request so a single slot buys several decisions, and they design for options they may never need, because a second visit costs another month. A change that needed one endpoint arrives with a configurable abstraction and an extension point that will still be there in five years.
Why it matters
Slow decisions do not only delay architecture; they systematically inflate it. Optionality is a rational hedge when the cost of returning is measured in weeks, so the organisation buys abstractions nobody needed with real engineering time, and pays again in every later change.
The metric also tells you whether governance is real. As latency rises, decisions start being taken outside the process and presented afterwards, or not at all. A forum whose request rate is falling may be succeeding or may have been abandoned, and the two look identical from inside it. The wait is also invisible in delivery metrics: work waiting for a decision is not in progress, so it is not late, and the lengthening lead time reads as engineering slowness.
Implementation patterns
- Instrument request and decision time. A row per decision with raised-at, decided-at, decider and reversibility class; a spreadsheet does for the first quarter.
- Report p50 and p90, never the mean. The tail is the behaviour that changes designs. A target worth holding is p90 under five working days; above two weeks, expect bundling and pre-emptive abstraction as standard.
- Publish the share of decisions taken outside the process alongside latency. Rising share is the clearest evidence the process has stopped being real.
- Delegate by blast radius, in writing. Reversible changes inside one team's boundary need no forum; shared contracts with more than two consumers, retention rules and one-way doors do. The requester must be able to apply the test without asking.
- Async review with a named reviewer and a deadline: a written proposal, one accountable reviewer, a 48-hour response commitment and default approval if the deadline passes. Default-approve is what converts a queue into a service level.
- Track the requests never raised. Ask teams quarterly which improvements they abandoned because approval was not worth the wait; that is the loss nothing else measures.
Industry example
Queueing behaviour makes the arithmetic unavoidable: with arrivals above service rate the backlog grows without bound, and wait time rises sharply as utilisation approaches capacity. Delivery research has repeatedly found approval wait dominating lead time, and it is visible in production wherever commit-to-release time has been decomposed and the waiting outweighed the working. The architectural twist is specific: the queue changes the content of what it serves, because requesters adapt their proposals to the cost of asking.
Failure scenarios
- Bundling, where unrelated decisions arrive together and the reviewer cannot assess any of them in the slot available.
- Pre-emptive over-design that outlives the forum by a decade.
- Shadow decisions, taken and presented retrospectively, so the record no longer describes the system, and abandoned small improvements — the silent loss of the changes that keep an estate tidy.
- A forum that adds attendees to improve quality, raising the cost per decision and lengthening the queue.
Trade-offs
Fast decisions are bought with accepted variance: delegated calls will sometimes be wrong and some inconsistency will appear across teams. Slow decisions buy consistency and pay in inflated designs and routed-around governance. The correct position is per decision, not per organisation: match the cost of the process to the cost of being wrong. Reversible and locally scoped — delegate and record. Irreversible, cross-company or regulated — queue it deliberately and say so.
When not to use it
Do not optimise this metric where the queue is correct. A choice that cannot be reversed for five years, a regulated change, a commitment to another company: those deserve a quorum, a record and a wait, and running them through a 48-hour default-approve creates a liability.
And do not measure it in a small organisation. With one team and one architect, latency is hours and the instrumentation is overhead. The metric earns its cost when decisions cross team boundaries often enough to form a queue, in practice above about five or six teams sharing contracts.
Interview question
Q: You find that the average time from raising an interface question to getting an answer is eighteen days, and that the designs arriving at review are more abstract than the requirements justify. Are those two facts related, and what would you change first?
What a strong answer covers: the causal link — optionality as a hedge against the cost of returning · measuring p90 and the share of decisions taken outside the process, not the average · delegation by reversibility first, because it removes demand rather than adding capacity · async review with default approval second · what deliberately stays in the queue · and the signal that it worked, namely smaller proposals rather than faster meetings.
Quick check
Quiz: A weekly forum serves six decisions and receives fourteen. Besides delay, what happens to the designs? They inflate: teams bundle unrelated changes to use one slot and add optionality they may not need, because a second visit costs another month.
Flashcard: What target holds architectural decision latency below the point where designs start to inflate? — p90 under five working days; above two weeks, expect bundling and pre-emptive abstraction.