A marketplace's infrastructure spend rises with growth. How should it distinguish healthy growth from deteriorating efficiency?
Show the full answer Hide the answer
Why total spend is uninformative
Total spend rising during growth is expected and tells you nothing about whether the architecture is improving or degrading. A business doubling its transactions and its spend has flat efficiency; one doubling transactions with a 50% spend increase has improved; one doubling transactions with a tripled spend is deteriorating — and all three look like "spend went up" on a monthly report.
What the unit metric must be
- Denominated in a business transaction, not in a technical unit. Cost per order, per active user, per seller, per delivery — something a non-engineer recognises and the business already tracks.
- Split by major cost category, because a single number is not actionable. Compute, storage, egress, third-party services and observability move for different reasons and are fixed by different people.
- Tracked over time as a trend, not compared to a benchmark. External benchmarks are almost useless because business models differ enormously in what a "transaction" costs.
- Broken down by customer segment or product line where the economics differ, since a blended number hides a product line that is unprofitable at the margin.
The uncomfortable finding it usually surfaces
Some segments or features cost more than they earn. A low-value order category whose fulfilment and support cost exceeds its margin, a free tier consuming a disproportionate share of infrastructure, a feature used by a small number of customers at high cost.
That is a product decision the metric enables and nothing else surfaces, and it is frequently more valuable than any engineering optimisation.
What makes the metric credible
Cost allocation that does not require tagging discipline. Tagging never survives contact with reality at scale, so the allocation should follow structural boundaries — account or project per workload, cluster per team — where attribution is automatic.
Teams that can see their own unit cost reduce it, and teams that cannot, do not. That is the entire mechanism, and the visibility is worth more than any specific optimisation it prompts.