A year into chargeback every product team's reported infrastructure cost is down between 15% and 25%. Company cloud spend is up 8%. Tag coverage is 100% and no commitment expired. Where is the money going, and what would you look at first?
Show the full answer Hide the answer
The first three things I would look at, and why in that order
- The shared and unallocated pool over time. If every allocated line fell and the total rose, the difference is by arithmetic in whatever is not charged back. This takes ten minutes and usually ends the investigation.
- Per-team consumption of internal platforms: log volume ingested, CI minutes, bytes scanned in the shared warehouse, topics and partitions held on the shared broker, internal API call counts. These resources are tagged to the platform team, so they are invisible in the team view that everyone is looking at.
- Where work moved rather than went away. Jobs pushed onto the shared batch cluster, functions moved into a managed service billed to a central account, chatty internal calls replacing local computation.
The diagnosis
Chargeback made each team's own resources expensive and left the shared platforms free at the point of use. Teams then behaved rationally and moved cost instead of removing it: more logging, because the pipeline is free to them; more CI, because the runners are free; heavier queries, because the warehouse is free; chattier services, because somebody else's mesh carries the traffic. Each local decision reduced a charged line and increased an uncharged one. A 20% average reduction against 8% net growth is the signature of exactly that, and it is not cheating by anyone. It is the incentive that was built.
The misleading signal
Tag coverage at 100% looks like the allocation problem is solved. Tags attribute resources to owners. They do not attribute consumption of a shared service to the consumer. A broker cluster can be perfectly tagged to the platform team while telling you nothing about who produced the traffic that grew 40% this year. Teams point at the coverage number, which is correct, and conclude the data is right, which it is not.
The fix
Meter consumption of shared services per consumer and publish it, even if the money never moves. Gigabytes ingested, partitions held, CI minutes, bytes scanned. The meter matters more than the charge, because behaviour responds to a number with your team's name on it. Set the meter before the platform becomes popular. Retrofitting one is a political exercise, because the first report names whoever grew fastest and the conversation becomes about the measurement.
Keep one rule: charge a team only for something it can change, and prefer a published meter to a disputed charge. Charging for an allocation a team cannot influence produces arguments and no savings. This failure mode has been a named capability in published cost-management guidance since about 2019, precisely because tagging alone never resolves it.
The alert that would have caught it earlier
A monthly ratio of allocated spend to total spend, alerting on a move of more than a few points in a quarter. Underneath it, spend per unit of business value at company level, which is the only number that cannot be improved by moving cost between accounts.
When this is the wrong answer
In an organisation with a dozen teams and one or two shared services the commons is small, and an 8% rise is more likely to be a price change, an expiring discount, or growth in a product line nobody attributed. Check the purchasing portfolio and the traffic before building a metering pipeline, because that pipeline is a permanent piece of infrastructure and it is expensive to build for the wrong reason.