Platform SLOs
Committing to reliability for internal consumers who cannot choose another provider.
5 to work through
-
intermediate
A platform team is about to publish its first commitments to internal consumers - 99.95% monthly availability on the ingress and the config service plus a 15-minute objective for deploys. What does the platform give up by publishing that, and when does the bill arrive?
3 min answer -
intermediate
Should an internal platform have SLOs, and what should they cover?
2 min answer -
intermediate
Why should a platform publish SLOs to internal consumers, and what should they cover?
2 min answer -
advanced
A product team commits to 99.95% availability. Their service depends on your platform's ingress, config service and secret manager. What do you tell them?
2 min answer -
advanced
Application teams say the platform is unreliable. The platform team's dashboard shows 99.95% on every component. How do you resolve this?
2 min answer
3 terms in this topic
Internal Consumer SLO
A reliability commitment made to teams who cannot switch supplier, which is why it must be measured from their side rather than from the platform's.
patternLong-Running Operation SLO
A reliability objective shaped for multi-minute platform operations like deploys and provisioning - share of runs reaching a terminal state within a …
metricPlatform SLO
A published reliability and performance objective the platform commits to for the teams that depend on it.
Neighbouring topics
Platform Engineering
General material on internal platforms as products with users, adoption and lifecycles.
Internal Developer Platform
The assembled surface teams actually touch, and what belongs behind it.
Paved Road & Golden Path
A supported default route that is easier than the alternatives rather than mandatory.
Self-Service Provisioning
Teams getting infrastructure without a ticket, and the guardrails that make that safe.
Service Templates
Scaffolding new services with observability, CI and security already wired in.
Platform APIs
Treating the platform's own interfaces as contracts with consumers and compatibility rules.
Platform Tenancy
Isolating teams sharing a cluster, account or pipeline fleet, and where isolation must be hard.
Cluster Architecture
How many clusters, split by what, and the blast radius each split buys.
Service Mesh Operations
What a mesh genuinely solves, its failure modes, and the cost of running one.
Container Image Strategy
Base images, layer hygiene, rebuild cadence, and patching a fleet of images.
Developer Environments
Local, remote and ephemeral environments, and the fidelity each can honestly claim.
Inner Loop & Outer Loop
Where an engineer's time actually goes, and which loop a platform investment shortens.
Abstraction Level Choice
How much to hide, and the leak that turns a helpful abstraction into a trap.
Platform Adoption
Migrating existing teams onto a platform without a mandate, and reading the adoption curve.
Platform Funding
Central cost, showback, chargeback, and justifying a team that ships no customer feature.
Platform API Deprecation
Removing something dozens of internal teams depend on, on a timeline that holds.
Guardrails vs Gates
Preventing a class of mistake automatically versus stopping to ask a human.
Platform Telemetry
Instrumenting the platform itself: usage, friction, and where teams leave the paved road.
Platform Team Topologies
Stream-aligned, enabling, complicated-subsystem and platform teams, and their interactions.