practice

Cost Incidence Mapping

also called Who-Pays Analysis, Payer And Horizon Columns

Adding a payer and a horizon to every row of a trade-off table, so that costs falling on teams who were not in the room stop being invisible to the team making the choice.

trade-offsoperating costownershiptechnology varietygovernance

A team compares two designs. Option A takes 5 weeks to build and introduces a new message broker. Option B takes 9 weeks and uses the broker the platform team already runs. The trade-off table has two columns, gains and costs, and on that table option A wins: it is 4 weeks cheaper and the "operational complexity" row is one line of prose.

The row is one line because the people who will pay it were not in the meeting. Somebody has to learn the new broker well enough to debug it at 03:00, test its backup path, patch it on the vendor's schedule, upgrade it twice a year and answer for it in the next audit. That work is continuous and it lands on a platform or on-call group with no vote in the decision.

Cost incidence mapping is the small discipline of making the table say who pays each row and over what period. It costs about 20 minutes per decision, because a cost that nobody has a name for is reliably estimated at zero.

Why it matters

Decentralised technology choice is defended on the grounds that the team closest to the problem decides best. That is true for the parts of the cost the team bears and it fails for the parts it does not, and the asymmetry has a shape: build cost is one-off, visible and borne by the chooser; operating cost is annual, invisible at decision time and usually borne by someone else. The result is an estate with more runtimes, datastores and brokers than anyone chose, each arriving through a locally rational decision.

The standing cost of one more component never stops: version upgrades, a backup path somebody must verify, enough on-call expertise to diagnose it under pressure, and a place in every security review. Across a 300-engineer organisation this routinely reaches a meaningful fraction of a platform engineer per component per year, which is why a platform group's headcount grows with variety rather than with traffic.

Implementation patterns

  • Two extra columns: payer and horizon. The payer column takes a team name, not "the organisation". The horizon takes "once", "annual" or "for the life of the system".
  • A rule that a row with no named payer is not costed. If nobody will claim it, the proposer owns it, and that is usually enough to change the proposal.
  • Convert annual rows into one number at the discount the business uses, so a 9-week build can be compared with 4 years of patching rather than argued against it.
  • Require the payer's acknowledgement for anything on the "for the life of the system" row — not approval rights over the design, just evidence that the bill was shown to the person receiving it.
  • Publish the estate's variety count: distinct datastores, brokers, runtimes and languages in production, each with a named owner. It is the clearest leading indicator of platform cost.

Industry example

A 300-engineer enterprise software vendor with more than 40 product lines runs a federated model: each product team picks its own storage. Over 6 years that produced 11 distinct datastores in production. No single choice was wrong, and each business case showed a saving against the central option. The bill arrived as a platform group whose entire capacity went to patching, backup verification and upgrade coordination, with none left for the developer-experience work the company needed to hit its revenue-per-engineer target. The fix was not a mandate but a column: new variety is approved where the choosing team carries its own on-call and upgrades, and refused where the cost lands on a group that was not in the room.

Failure scenarios

  • The platform team becomes the payer of last resort and its roadmap is consumed by other teams' choices, which looks like poor platform prioritisation and is not.
  • Reorganisation transfers an unpriced cost: the team that chose the component is dissolved and the component becomes nobody's, where it stays until it breaks.
  • A build-versus-buy case that counts licence fees and not operating labour, so self-hosting always wins on paper and loses in the second year.
  • Theatre: the columns get filled with "platform" on every row and nobody checks, which creates a record of consent that was never given.

Trade-offs

Choose Gains Pays
Incidence mapping on every significant choice Total cost visible before commitment; variety grows deliberately 20 minutes per decision and some genuinely slower local choices
Local autonomy with no incidence column Fast decisions and high team ownership Operating cost accumulates where nobody chose it

The deeper cost is political. Naming a payer makes explicit a conflict that was previously resolved by the platform group quietly absorbing work, so the first few mapped decisions will be slower and more contested. That contest is the point: it is the decision being made by the people who pay for it.

When not to use it

Do not apply it to reversible, short-lived or self-contained choices. A library inside one service, a scheduled job one team can delete, a prototype with a written kill date: the horizon column reads "once" and the exercise is overhead.

It is also wrong where the choosing team genuinely is the payer — its own on-call rota, its own infrastructure budget, the authority to run what it picks. Then the incidence is already internalised. The practice earns its cost where the deciding boundary and the paying boundary differ, which above about 100 engineers is most of the time.

Interview question

Q: A product team's business case shows self-hosting a search engine saves 180k dollars a year against the managed service. The case is arithmetically correct. What would you add before approving it, and what would make you approve it anyway?

What a strong answer covers: naming the payer and horizon of each cost row · the labour the case omits — upgrades, backup verification, capacity work, on-call expertise, security review — and that it is annual rather than one-off · whether that labour sits in the proposing team's own budget and rota · the conditions under which self-hosting still wins, namely that the team carries its own operation, the workload is large enough that the saving exceeds a loaded engineer's cost, and the company already runs the component elsewhere · and the signal to watch afterwards, whether the platform group's unplanned work rises in the two quarters after go-live.

Quick check

Quiz: Your trade-off table lists gains and costs for each option. Which column is missing and what does it change? Who pays each row and over what horizon — adding it reveals that the fast-to-build option transfers a permanent annual cost to a team with no vote, which frequently reverses the ranking.

Flashcard: Why does decentralised technology choice drift toward high variety even when every individual decision is sound? — Build cost is one-off and borne by the chooser; operating cost is annual and borne by someone else, so it is estimated at zero.