advanced 3 min answer

A review board rejected a team's design in week 3 and required it to use the shared notification service instead. 14 months later that shared service's queue backed up for 4 hours during a marketing send; payment confirmations went with it, because they travelled the same path. The board's decision record names the mandate, the date and the attendees, and nothing else. What failed, and which governance decision made the outage possible?

architecture-review-boardshared-servicesaccountabilityblast-radiusdecision-records
Show the full answer Hide the answer

The trigger

A marketing campaign pushed a large batch of low-priority messages into a queue that also carried payment confirmations. The queue had one class of service, so the confirmations waited behind the campaign, and a 4-hour backlog on notifications became a 4-hour payments incident.

The proximate fix is obvious: separate priority classes, or separate queues. The interesting question is why a team that would have built its own notification path — and would have been paged for it — ended up on a shared dependency whose failure behaviour nobody had specified.

Why the mandate propagated the failure

A review board that mandates a dependency moves the risk without moving the accountability. Three specific gaps do the work:

  • No service level attached to the mandate. The team was required to depend on the shared service, but nothing said what latency or availability it could rely on. With no number, the team could not have designed a fallback even if it had wanted to, because there was nothing to design against.
  • No owner named in the decision. The shared service had a maintaining team, which is not the same as a team accountable for the consequences of every mandated consumer. When the incident happened, the payments team was paged for a system it had been told to use and did not control.
  • No exit condition. A mandate with no review date is permanent. Fourteen months later the original reasoning — reuse, one integration with the SMS vendor — was never re-tested against the fact that the service now carried revenue-critical traffic.

Why detection lagged

The notification service's own dashboards were healthy in the sense that mattered to its owners: throughput was high and nothing errored. Queue depth was rising and no alert distinguished whose messages were in it. Payment confirmation latency was measured by the payments team, which had no reason to watch a notification queue. The signal existed in two places and was correlated in neither.

The structural fix versus the tempting local fix

The tempting fix is priority queues in the shared service. Do it — it takes a day and removes this failure. The structural fix is that a mandate becomes a contract or it does not get issued: every board decision that requires a team to depend on something carries a stated service level, a named accountable owner, and a review date. The board then discovers how few mandates it is willing to issue on those terms, which is the point.

When this is the wrong answer

Where the shared component is genuinely undifferentiated and its failure is not customer-visible — an internal metrics agent, a logging sidecar — demanding a full contract for every mandate is bureaucracy that buys nothing. The line is whether the mandated dependency sits on a revenue or safety path. If it does, the board is making a reliability decision, not a reuse decision, and reliability decisions need numbers.

What a strong answer adds

That the board's incentive is inverted. It is measured on reuse and consistency, and the cost of a bad mandate lands 14 months later on someone else's pager. Publishing the incidents traced to mandated dependencies is the only feedback loop that corrects it, and most boards have never seen that list because nobody assembles it.