An interviewer tells you: you now own architecture review for 400 engineers in 40 teams. Median time from submitting a design to getting a decision is 19 days, and three teams shipped significant designs without review last quarter. Talk me through what you change.
Show the full answer Hide the answer
What the interviewer is testing
Whether you treat the review as a queue with a service rate or as a cultural problem. Candidates who answer "set expectations and build relationships" have not noticed that 19 days is arithmetic, and that the three teams who skipped it were behaving rationally.
The clarifying questions that change the answer
- Is the review advisory or a gate? If it can block a release, the queue is a dependency on the critical path and the fix is different from an advisory board.
- Is the 19 days waiting or rework? Waiting is capacity; rework is an unclear bar. The remedy differs.
- What did the three skipped designs do? This is the evidence, not the violation.
The arithmetic to do out loud
40 teams producing roughly one significant design a month is about 10 submissions a week. A review honestly costs 4 reviewer-hours — two reading and preparing, one in session, one writing the decision — so 10 a week needs about 40 reviewer-hours a week, which is one dedicated reviewer running flat out.
Queues punish high utilisation. At 90% utilisation, waiting time is roughly nine times the service time; at 70% it is about 2.3 times. That accounts for 19 days with nobody being slow, because the queue — not the reviewer — is where the time goes. Two levers exist: raise capacity or reduce arrivals. Only one of them is free, and a review process fails quietly in the second way: teams route around it and the organisation stops knowing which designs went to production unexamined.
A strong answer's arc
- Triage on reversibility and blast radius. Review only designs that are hard to undo or that cross team boundaries: new persistence of personal data, tenant isolation, money movement, a new runtime or language, a shared schema. In most organisations that is 20 to 30% of submissions, which takes arrivals to 2 or 3 a week and utilisation into the range where the queue drains.
- Everything else gets a checklist and a named reviewer inside the team, with the design still written down. The point is not less writing, it is less centralised waiting.
- Publish the triage predicate so a team can tell before submitting which path they are on. Unpredictable gates are the thing people route around.
- Change the measure. Stop counting reviews held. Track decision latency p50 and p90, the share of designs reviewed before implementation starts, and findings that changed the design per review. If that last number is under about one, the review is theatre and should be narrowed further.
- Go and read the three skipped designs, all of which are now in production. Compute what review would have caught. If the answer is nothing, the gate was mis-scoped and you have just found your triage rule.
The decision rule to state plainly: review a design when undoing it would take more than a sprint or would require another team's cooperation; otherwise prefer a checklist and a team-local reviewer.
Common weak answers
- "Add reviewers." It works until demand grows, costs senior headcount linearly, and leaves the real defect — every design queueing behind every other design — untouched.
- "Make it fully asynchronous." Written review is good and it is where disagreement goes quiet. Keep a short synchronous slot for the designs that are genuinely contested.
- "Enforce it through the pipeline." Blocking a deploy on a review stamp raises compliance and hides the cost in release latency.
What a strong answer adds
Review is capacity, so it should be spent where a decision is expensive to reverse. Say what you will stop reviewing, name who holds the decision when the board disagrees with a team, and commit to a number: for example p90 decision latency under 5 working days within a quarter, published monthly.