FinOps Practice
The operating model that makes cost a continuous engineering concern rather than a periodic finance exercise.
Definition
FinOps is the practice of bringing financial accountability to variable cloud spend: engineering, finance and product share responsibility for cost, informed by timely and accurate data.
The three phases, applied continuously
Inform. Accurate allocation, visibility for the people who can act, forecasting, and unit economics — cost per order, per user, per transaction. Nothing else works without this, which is why tagging enforcement is the first move in any remediation.
Optimise. Rightsizing, commitments, waste elimination, architectural change. Ordered by return per unit of effort, not by whichever is most technically interesting.
Operate. Cost as a standing part of engineering practice: budgets and targets, anomaly alerting, cost considered in design review, and someone accountable.
What separates a working practice from a cost-cutting project
- Continuous, not periodic. A quarterly cost review produces a spike of activity followed by nine months of drift.
- Engineers see their own numbers, in a form they can act on, next to their unit cost. Reports that go only to finance change nothing.
- Cost is a design consideration, estimated in review alongside latency and availability — so it is decided before it is spent rather than remediated afterwards.
- Anomaly detection, so a runaway job or a mistaken configuration is caught in a day rather than at month end.
- Unit economics as the primary metric. Total spend should rise with a growing business; unit cost should fall.
What it must not become
A veto function. A practice that blocks engineering decisions on cost grounds without weighing them against delivery speed and reliability will be routed around, and the shadow infrastructure that results is both more expensive and invisible.
The correct posture is making the cost of options visible at the moment of choice, and letting the team that owns the outcome decide.
Failure scenarios
- Cost reviews without allocation, so the discussion is about unattributable totals.
- A central team optimising other teams' workloads without context, and being correctly resisted.
- Cost targets without reliability targets, producing outages that cost more than the savings.
- Data reported monthly in arrears, too late to connect a change to its effect.
- One-off cost projects with no mechanism to prevent recurrence.
Interview question
"How would you make cost a routine engineering concern without slowing delivery?"