Cluster Architecture
How many clusters, split by what, and the blast radius each split buys.
6 to work through
-
advanced
A Kubernetes platform shows GPU utilisation below 40% because jobs request whole nodes. How should bin packing, gang scheduling, preemption, partitioning and queue-based admission raise utilisation without harming priority jobs?
3 min answer -
advanced
A Kubernetes upgrade to 1.24 removes a deprecated node label. Within minutes the cluster's pod network loses all routes and stays down for over five hours, and the configuration responsible is not in any repository. What kind of dependency is this, and how would you have found it beforehand?
3 min answer -
advanced
A platform must decide between one large cluster and many small ones. What are the trade-offs?
2 min answer -
advanced
How many Kubernetes clusters should an organisation run, and what decides the boundaries?
2 min answer -
advanced
How should compute clusters be structured for a platform running many teams' workloads with very different profiles?
2 min answer -
advanced
Your platform team of five spends nearly all its capacity upgrading control planes across 38 Kubernetes clusters. How did this happen and how do you fix it?
2 min answer
4 terms in this topic
Cloudflare: Every Server Runs Every Service
Rather than dedicating machines to roles, Cloudflare runs the full software stack on every server in every location, which turns capacity into a sing…
practiceCluster Architecture
How many clusters, how they are divided, and what shares fate — decisions about blast radius and operational burden rather than about capacity.
conceptCluster Split Axis
The dimension along which the estate is divided into clusters — the single decision that determines both blast radius and operational load.
conceptSelector Dependency
A production dependency that exists only as a string matching a label or annotation, so it is invisible to code search and to release notes, and a pl…
1 artifact you would hand over
Neighbouring topics
Platform Engineering
General material on internal platforms as products with users, adoption and lifecycles.
Internal Developer Platform
The assembled surface teams actually touch, and what belongs behind it.
Paved Road & Golden Path
A supported default route that is easier than the alternatives rather than mandatory.
Self-Service Provisioning
Teams getting infrastructure without a ticket, and the guardrails that make that safe.
Service Templates
Scaffolding new services with observability, CI and security already wired in.
Platform APIs
Treating the platform's own interfaces as contracts with consumers and compatibility rules.
Platform Tenancy
Isolating teams sharing a cluster, account or pipeline fleet, and where isolation must be hard.
Service Mesh Operations
What a mesh genuinely solves, its failure modes, and the cost of running one.
Container Image Strategy
Base images, layer hygiene, rebuild cadence, and patching a fleet of images.
Developer Environments
Local, remote and ephemeral environments, and the fidelity each can honestly claim.
Inner Loop & Outer Loop
Where an engineer's time actually goes, and which loop a platform investment shortens.
Abstraction Level Choice
How much to hide, and the leak that turns a helpful abstraction into a trap.
Platform SLOs
Committing to reliability for internal consumers who cannot choose another provider.
Platform Adoption
Migrating existing teams onto a platform without a mandate, and reading the adoption curve.
Platform Funding
Central cost, showback, chargeback, and justifying a team that ships no customer feature.
Platform API Deprecation
Removing something dozens of internal teams depend on, on a timeline that holds.
Guardrails vs Gates
Preventing a class of mistake automatically versus stopping to ask a human.
Platform Telemetry
Instrumenting the platform itself: usage, friction, and where teams leave the paved road.
Platform Team Topologies
Stream-aligned, enabling, complicated-subsystem and platform teams, and their interactions.