Two hundred control planes per cluster: ten years of SAP's Gardener
How SAP's Gardener changed between 2017 and 2026 while running tenant Kubernetes control planes as ordinary workload on five infrastructures: what it adopted, what it withdrew, what each decision cost, and what the shared-host model charges in blast radius.
A fleet platform has to give hundreds of teams their own cluster, on whichever infrastructure each product sells on, cheaply enough that asking for one is not a budget conversation. Gardener is the rare case where a large company's answer can be read rather than guessed at, and this guide reconstructs it from thirty-five enhancement proposals, a feature gate table that dates every mechanism adopted and withdrawn, 759 timestamped releases, two complete release archives compared across six years, and the three 2025 advisories that price the architecture's one structural risk. A reader finishes able to choose a shape for a fleet control plane, defend a packing ratio and a per-tenant failure domain in review, and name the two failure classes the model creates.
The parts of the platform that survived a decade are the ones with a one-to-one Kubernetes analogue, and the two mechanisms deleted without ever reaching general availability are one adopted upstream component used against the grain (the Konnectivity tunnel, alpha in release 1.6, gone in 1.27) and one in-house invention that fused two upstream controllers into one (the HVPA autoscaler, alpha since release 0.31 in 2019, removed in 1.109 in November 2024).
What you get out of it
- Tenant control planes run as pods in a shared host cluster at a documented default of 200 per host, which converts cost that would scale with cluster count into cost that scales with load; the comparator is a published $0.10 per cluster per hour.
- Packing tenants onto a shared host makes the tenancy boundary a validation function: all three 2025 advisories end with a project administrator obtaining control over the host cluster, and one of them had to be fixed in four provider extensions separately.
- Availability was deliberately a scheduling property rather than a replica count until GEP-20 in August 2022 admitted that single-replica etcd made a zone failure a tenant outage, after which the failure domain became a per-tenant field.
- Any controller that acts on the absence of a heartbeat needs a second opinion about the heartbeat path: the dependency watchdog exists because node replacement destroyed healthy machines during a network fault, and it scales the replacing controller to zero instead.
- A fork of a dependency is a dated loan: sitting on a fork of a relicensed log engine bought security updates only, and the escape was to standardise collection first so the store became replaceable.
Scope
Why this, now. The fleet pattern is spreading from Kubernetes platform teams to anyone running per-tenant control planes, and Gardener's 2025 and 2026 record (three host-takeover advisories, in-place node updates driven by GPU scarcity, a forked log store finally replaced) is the freshest public evidence of what the model costs to operate.
What it does not cover. SAP's wider estate (HANA, ABAP, the Business Technology Platform application layer), the Gardener dashboard and CLI surface, the commercial history, and SAP's own operated scale and spend: this session's network policy reached only raw.githubusercontent.com, the Go module proxy, pkg.go.dev and one vendor price list, and refused every engineering blog, talk, paper archive, status page and GitHub's own issue and pull request API, so there is no blog, talk or paper evidence here and no published incident review.
Other field guides
The parts that outlived the product: ten years of Docker, read from its own repositories
A decade of one company's architecture told through the sequence every platform team eventually faces: bundle to win the workflow, extract components…
38 sources · 9 organisations · 5 postmortemsBoundaries without a network hop: ten years of Shopify, read from its own artefacts
Reconstructs a decade of one company's architecture from artefacts rather than announcements: repository archive notices, release notes of the langua…
30 sources · 8 organisations · 2 postmortemsYour singletons choose where you fail: ten years of GitLab's architecture
A decade of one company's platform evolution, reconstructed entirely from primary artefacts it publishes in Git repositories: 207 design documents, t…
28 sources · 2 organisations · 5 postmortems