intermediate 3 min answer Multiple choice

A platform team of 9 serving 300 engineers changes model. It stops publishing modules and libraries that teams run themselves and starts operating the runtime: it owns the clusters, the pipelines and the base images, and it holds the pager for them. Within two quarters the share of services on the supported path goes from 60% to 95% and a runtime upgrade that used to take three quarters takes five weeks. What has the organisation given up?

platform-teamsteam-topologieson-callcognitive-loadownership
Pick one
Show the full answer Hide the answer

What is gained

Ownership of the runtime buys the ability to change the estate without asking 300 people. That is the whole case, and the numbers in the stem are what it looks like: 95% on the supported path, a runtime upgrade in five weeks instead of three quarters. An enablement platform can never produce that, because every change is a request to 40 backlogs with 40 different priorities. When the change is not optional — a runtime leaving support, a CVE, a control an auditor requires — the difference between five weeks and three quarters is the whole argument.

What is paid

Every incident now starts at the platform team's door, because the platform is the only component present in all of them. A service times out: is it the service, the mesh, the node pool, a base image change, or the pipeline that shipped it? Somebody has to answer that, and the team that owns the substrate is the only team that can. Triage load scales with the number of services, not with the number of platform components.

The arithmetic is unforgiving. Nine engineers, 300 served: an on-call rotation of nine is roughly one week in nine per person, and each week carries the triage volume of 300 services' worth of ambiguity. At 10 to 15 interrupts a week the rotation is a job; past that it consumes the roadmap, and the platform stops shipping the improvements that justified the model. This is the cognitive-load argument in Team Topologies running in reverse: load removed from stream-aligned teams did not vanish, it moved.

When the bill arrives

Not at the switch. It arrives when adoption is nearly complete, which is also when the model looks most successful. At 60% adoption the unsupported 40% absorb their own incidents. At 95% there is nowhere else for an incident to go, so the interrupt rate roughly triples over the same two quarters that the adoption metric is being celebrated.

How to keep the option to reverse

  • Publish a platform contract with numbers — what the runtime guarantees, what it does not, and which failures are the service's own. Without it every symptom is arguably the platform's.
  • Fund the rotation explicitly. If ownership is worth five-week upgrades, it is worth 2 to 3 of the 9 engineers being unavailable for project work, budgeted rather than discovered.
  • Keep a first-line triage path that is not the platform team — a runbook plus a small rota drawn from stream-aligned teams answers the "is it me or the platform" question in most cases.
  • Own the substrate and not the application runtime semantics. Owning the cluster is finite; owning "whatever your service does on the cluster" is not.

Why the other options fail

"Lost the ability to upgrade centrally" inverts the outcome: central ownership is precisely what produced the five-week upgrade. This is the cost of the model the team left behind.

"Infrastructure cost" usually moves the other way — consolidation onto shared node pools improves utilisation. The expensive resource here is the nine people, not the compute.

"Very little since the headcount is the same" is the answer that gets platform teams burned out. The headcount may be the same; the work is not. Enablement work is schedulable and ownership work is interrupt-driven, and the two do not fit in one person's week.

When ownership is the wrong call

When the variety is legitimate. If four product lines have genuinely different runtime needs — a latency-critical service, a batch estate, a mobile backend, an ML platform — one owned runtime becomes a lowest-common-denominator compromise that each of them works around. Prefer enablement when the estate's diversity is a product fact rather than an accident, and prefer ownership when the diversity is history nobody is defending.