Platform Tenancy
Isolating teams sharing a cluster, account or pipeline fleet, and where isolation must be hard.
6 to work through
-
advanced
A platform serves internal teams with very different scale and criticality. Should they share infrastructure?
1 min answer -
advanced
An internal platform hosts workloads from many teams. What isolation model should it use, and how are noisy neighbours handled?
2 min answer -
advanced
An internal platform serves many product teams with very different scale and criticality. How should isolation between them be designed?
2 min answer -
advanced
On 11 December 2024 OpenAI deployed a new telemetry service into its Kubernetes clusters. Its configuration caused every node to perform Kubernetes API operations whose cost scaled with the size of the cluster; API servers saturated and DNS-based service discovery failed across most large clusters, with services unavailable from 15:16 to 19:38 PST. The change had been tested in staging. What structurally failed, what could staging not test, and where would copying their response be a mistake?
3 min answer -
advanced
On a shared cluster serving twelve teams, one team's new operator creates a custom resource per user session, reaching about 50,000 objects in a day. Other teams start seeing slow deploys and intermittent API timeouts. What is happening inside the cluster?
3 min answer -
advanced Multiple choice
Twelve teams, forty services. One shared Kubernetes cluster or a cluster per team? Justify.
1 min answer
4 terms in this topic
Cluster Multi-Tenancy
Multiple teams or workloads sharing one cluster, isolated by namespaces, quotas and policy rather than by separate infrastructure.
patternNamespace Tenancy
Isolating teams within a shared cluster by namespace plus quota plus network policy, and being explicit about what that does not isolate.
conceptNoisy Neighbour
One workload degrading others by consuming a shared resource - and specifically the resources nobody set limits on, such as control-plane capacity, n…
practiceObject Count Quota
A per-tenant limit on how many control-plane objects a team may create, which is the control that prevents one team's design choice from exhausting t…
Neighbouring topics
Platform Engineering
General material on internal platforms as products with users, adoption and lifecycles.
Internal Developer Platform
The assembled surface teams actually touch, and what belongs behind it.
Paved Road & Golden Path
A supported default route that is easier than the alternatives rather than mandatory.
Self-Service Provisioning
Teams getting infrastructure without a ticket, and the guardrails that make that safe.
Service Templates
Scaffolding new services with observability, CI and security already wired in.
Platform APIs
Treating the platform's own interfaces as contracts with consumers and compatibility rules.
Cluster Architecture
How many clusters, split by what, and the blast radius each split buys.
Service Mesh Operations
What a mesh genuinely solves, its failure modes, and the cost of running one.
Container Image Strategy
Base images, layer hygiene, rebuild cadence, and patching a fleet of images.
Developer Environments
Local, remote and ephemeral environments, and the fidelity each can honestly claim.
Inner Loop & Outer Loop
Where an engineer's time actually goes, and which loop a platform investment shortens.
Abstraction Level Choice
How much to hide, and the leak that turns a helpful abstraction into a trap.
Platform SLOs
Committing to reliability for internal consumers who cannot choose another provider.
Platform Adoption
Migrating existing teams onto a platform without a mandate, and reading the adoption curve.
Platform Funding
Central cost, showback, chargeback, and justifying a team that ships no customer feature.
Platform API Deprecation
Removing something dozens of internal teams depend on, on a timeline that holds.
Guardrails vs Gates
Preventing a class of mistake automatically versus stopping to ask a human.
Platform Telemetry
Instrumenting the platform itself: usage, friction, and where teams leave the paved road.
Platform Team Topologies
Stream-aligned, enabling, complicated-subsystem and platform teams, and their interactions.