Shared-Service Commons
also called Free-at-Point-of-Use Platform, Uncharged Internal Platform
The failure mode where cost allocation prices each team's own resources but leaves internal platforms free to use, so teams move cost onto the shared estate instead of removing it and total spend rises while every team's line falls.
A year into a chargeback programme, every product team's reported infrastructure cost is down between 15% and 25%. Company cloud spend is up 8%. Tag coverage is 100% and no commitment expired. Nobody cheated. The incentive that was built is the one that operated.
Chargeback made resources a team owns expensive and left the log pipeline, the CI fleet, the service mesh, the shared broker and the analytics warehouse free at the point of use. Teams responded rationally: log more, because the pipeline costs them nothing; run more CI, because runners are free; push batch work onto the shared cluster; query the warehouse harder; let a chatty internal call replace a local computation. Each decision reduced a charged line and increased an uncharged one.
Why it matters
It produces a cost programme that reports success while the company's spend grows, which destroys the credibility of cost work for years afterwards. It also distorts architecture. A team choosing between caching locally and calling a shared service is making a technical decision under a price signal that says the shared service is free, so the estate drifts toward designs that are cheap for teams and expensive for the company.
The second effect is on the platform team, which absorbs growth it did not cause, cannot forecast, and has no mechanism to push back on. Being unable to attribute demand is the same as being unable to refuse it.
Implementation patterns
- Meter consumption per consumer for every shared service, on a unit the consumer can influence: gigabytes ingested, partitions held, CI minutes, bytes scanned, requests served.
- Publish the meter before charging for it. The number with a team's name on it drives most of the behaviour change; moving money is a separate and more political decision.
- Install the meter before the platform becomes popular. Retrofitting one onto a busy platform is a negotiation, because the first report names whoever grew fastest.
- Charge only for what a team can change. An allocation a team cannot influence produces arguments and no savings.
- Track the ratio of allocated to total spend monthly and alert on a move of more than a few points in a quarter. This is the cheapest detector for the whole failure mode.
- Keep one company-level unit metric (spend per order, per active user, per booking) that cannot be improved by moving cost between accounts.
Industry example
The dynamic is the commons problem described by Garrett Hardin in 1968, applied to internal platforms, and it is why the FinOps Foundation's published guidance from 2019 onward treats allocation of shared and common costs as a distinct capability rather than a detail of tagging. Platform teams at organisations running internal developer platforms report the same sequence: the platform is free while adoption is the goal, adoption succeeds, and the platform's own bill becomes the fastest-growing line in the company with no way to attribute it.
Failure scenarios
- Team lines fall, total spend rises, and the cost programme is declared a failure by finance and a success by engineering.
- Tag coverage is used as proof that allocation is correct. Tags attribute resources, not consumption of a shared service, and the difference is invisible in a coverage metric.
- The platform team is given a budget cut for growth caused by its consumers, and responds by degrading the platform.
- A late retrofit of metering turns into a quarter of disputes about the measurement rather than the spend.
- Charging for an unchangeable allocation, such as splitting a mesh bill by headcount, which teams correctly ignore.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Showback on shared services | Behaviour change with low political cost | No budget consequence for a team that ignores it |
| Full chargeback of shared services | Strongest signal and honest team budgets | Metering infrastructure plus a real risk of discouraging use of the platform you want adopted |
| Leave shared services free | Maximum adoption of the platform | Cost migrates into the commons and becomes unattributable |
When not to use it
With a dozen teams and one or two shared services the commons is small, and building a metering pipeline costs more than it can return. Check the purchasing portfolio, a price change and product growth first, since these explain most unexpected increases at that size. Be careful too when a platform is early and adoption is the goal: pricing a platform before anyone uses it suppresses the adoption that justified building it. The right moment is when the platform's own spend becomes material, which is usually well before anyone has planned for it.
Interview question
Q: Every team's cloud cost is down 20% after a year of chargeback and company spend is up 8%. Tag coverage is complete. What happened, and what would you change?
What a strong answer covers: identifying the unallocated pool by arithmetic before investigating anything else; explaining that tags attribute resources and not consumption; naming the specific shared services that absorb migrated cost; proposing per-consumer meters with showback first; and the ratio alert that would have caught it in the first quarter.
Quick check
Quiz: Under chargeback, why does log volume tend to grow? The pipeline is free at the point of use for the producing team, so reducing local resources and logging more improves that team's charged line.
Flashcard: What do tags not attribute? Consumption of shared services. A broker cluster tagged perfectly to the platform team says nothing about who produced its traffic.