intermediate 3 min answer

Review this cost programme - a monthly per-team spend report circulated as a PDF, a quarterly 5% efficiency target for every team, an enforced tagging policy, and one analyst who files optimisation tickets. Tag coverage is 97% and cost per request has not moved in a year. What would you remove, what would you change, and what would you leave alone?

showbackfeedback-latencydesign-reviewefficiency-targetspull-request
Show the full answer Hide the answer

What is actually required

A cost programme has one job: make the person who can change the spend aware of it while they can still change it cheaply. Everything in this programme is about producing accurate numbers, and nothing in it is about timing.

The defect

Count the delay. A resource is created in a pull request. Billing data for it appears hours to days later, because providers meter and consolidate usage on their own schedule rather than yours. The monthly report lands up to four weeks after that. By the time the number reaches the engineer, three to ten weeks have passed, the change has shipped, and the cheap moment to alter it is gone. The programme is not short of information; it is short of feedback latency. Accuracy was improved from 90% to 97% tag coverage while the thing that blocks action was untouched.

What I would remove, and why it is safe to

  • The quarterly blanket 5% target. A percentage applied to every team regardless of its shape rewards whoever has the most slack and punishes whoever already runs lean. Worse, the easiest 5% is always redundancy: the standby replica, the spare zone capacity, the retained backup. Blanket targets buy savings out of the reliability budget without anyone deciding to, and the failure arrives as an outage in a quarter nobody connects to the target that caused it.
  • The PDF. Not the data, the format. A number nobody can query is a number nobody will check.

The one change that matters

Put a cost delta in the change path. Price the resources a change creates from the infrastructure plan against the rate card, and post it on the pull request and in the design review. A plan-time estimate is accurate to perhaps plus or minus 30% and arrives weeks earlier than an exact figure, and for the decisions that matter that trade is strongly in favour of speed. "This adds a managed gateway and a load balancer, roughly $90 a month each, times every environment" changes a design while it is still a diff.

Pair it with one threshold: any change whose estimate exceeds a stated figure needs a named approver. That is the only part that needs governance, and it is the part that survives contact with a real team in production.

What I would leave, even though it looks odd

  • The tagging policy, enforced at creation. It is unglamorous and it is the prerequisite for everything else. Keep the enforcement, drop the backfill campaigns.
  • The analyst. Move them out of the ticket queue and into design reviews and commitment negotiations, where one decision is worth more than a hundred tickets.
  • The monthly report, as a queryable dataset, for finance. It is the wrong tool for engineers and the right one for forecasting.

When this is the wrong answer

Below roughly $50k a month of total spend, do none of this. Building and maintaining plan-time cost estimation is on the order of two engineer-weeks plus upkeep, call it $15k in the first year. Against a $600k annual bill that is justified by a 3% improvement. Against a $60k annual bill it never pays back, and the right choice is a monthly glance at the top five line items. The rule flips when a single team can create spend faster than a month of reporting can reveal it, which is true of anything with an autoscaler attached.

Common weak answers

  • "Move from showback to chargeback." Changing who holds the budget does not change when they find out. The latency defect survives the transition intact.
  • "Improve tag coverage to 100%." The remaining 3% is shared infrastructure that tags cannot attribute anyway.