intermediate 3 min answer

A platform team is about to publish its first commitments to internal consumers - 99.95% monthly availability on the ingress and the config service plus a 15-minute objective for deploys. What does the platform give up by publishing that, and when does the bill arrive?

platform-sloserror-budgetsdependenciesinternal consumerscommitments
Show the full answer Hide the answer

What is gained

Consumers stop guessing. A team with a 99.9% product target can finally do the arithmetic: four platform dependencies in series at 99.95% each multiply out to about 99.8%, which is already worse than their target, so either they add a fallback, the platform removes a dependency from the path, or the product target changes. That conversation cannot happen without published numbers, and in its absence every team either over-builds its own fallbacks or discovers the gap during an incident.

What is paid

The platform's own change velocity now competes with its promise. 99.95% of a month is about 21.6 minutes of unavailability, total. A single mesh upgrade that restarts the fleet, one config-service failover, or one bad rollout can spend most of a month's budget in one afternoon. The platform must now stage its own upgrades, keep an n-1 path alive, and in practice schedule less of its own work per month than before.

You inherit your dependencies' numbers as your floor. If the managed control plane you build on is itself offered at 99.95%, you cannot promise 99.95% on top of it. The published figure has to be lower, or the architecture has to take that dependency off the serving path, which is real engineering, not a wording change.

You now own a measurement system. An internal consumer cannot switch supplier, so a number measured from the platform's own side is worthless to them; it has to be measured from theirs. That means synthetic probes running as a tenant, per-consumer success rates, and someone maintaining it.

And the deploy objective is a different kind of claim. A 15-minute deploy is a long-running operation, not a request. Availability percentages say nothing useful about it; the honest form is the share of deploys reaching a terminal state within 15 minutes with no manual repair, with stuck deploys counted separately from failed ones. Publish the wrong shape and you will report 100% availability on a control plane whose deploys hang.

When the cost becomes visible

At the first upgrade-heavy quarter, when the platform has to choose between a version bump and its budget. And at the first regional dependency incident, when consumers ask whether the cloud provider's outage counts against your number. Decide that in advance and write it down, because deciding it afterwards is how a published objective stops being believed.

How to keep the option to reverse

Publish an objective with a review date rather than a contract. Start with the two components you can already measure from the consumer side and say explicitly that the rest are unmeasured. Pair each number with an error-budget policy that states what the platform stops doing when the budget is gone - a number with no policy attached changes nobody's behaviour, and removing it later costs more trust than never publishing it.

When not to publish one

Below roughly ten consumer teams, a status page, a named support rota and a response-time promise deliver most of the value for none of the overhead. An SLO is a mechanism for coordinating at a scale where you cannot talk to every consumer. If you can, talk to them.