Fleet control planes  / field guide
Practitioner field guide · 1 October 2026

Two hundred control planes per cluster: ten years of SAP's Gardener

Gardener is the system SAP started in 2017 to give its own teams Kubernetes clusters on five infrastructures, and nearly every architectural decision it took is public: thirty-five enhancement proposals, a feature gate table that dates every mechanism it adopted and every one it withdrew, 759 released versions with timestamps, and three 2025 advisories that price the architecture's one structural risk. This guide reconstructs what changed between 2019 and 2026 and why, so that a reader about to build or buy a fleet control plane can argue the same trade-offs with numbers instead of taste.

27 primary sources 35 decision records 3 security advisories Evidence through September 2026 Read: 19 min
01

The territory

One team has to hand hundreds of other teams their own compute environment, on whichever infrastructure each product is sold on, and keep the unit cost low enough that asking for one is not a budget conversation.

200
Tenant control planes one host cluster is configured to hold, by default
152
Minor releases since February 2020, one every 16 days
$0.10
Per cluster per hour, what Google charges for the same service
3
2025 advisories where a tenant admin could take over the host cluster

Three groups have solved this in production. The hyperscalers solved it and published a price list rather than a design: Google's flat fee of "$0.10 per cluster per hour (charged in 1 second increments) applies to all GKE clusters irrespective of the mode of operation, cluster size, or topology", which is roughly $73 a month for a control plane you never see. The Kubernetes community solved a narrower problem with Cluster API, which Gardener's own documentation characterises precisely: "Cluster API primarily harmonizes how to get to clusters, while Gardener goes a significant step further by also harmonizing the clusters themselves." And SAP solved it for itself, starting in 2017 according to the copyright line that still ships in the repository, and then wrote the decisions down in public.

That last part is why this guide exists. Gardener is the rare case where a large company's fleet platform can be read rather than guessed at: the enhancement proposals state the motivation, the goals, the non-goals and the rejected alternatives; the feature gate table records when each mechanism entered and when it was taken out again; the module proxy timestamps every release; and the vulnerability record states, in the vendor's own words, what a tenant could reach. Set against that, the public record of how AWS, Google, Microsoft and the other managed-Kubernetes operators run their own fleets is a marketing page and a status history.

The sharpened question for this guide: what does a decade of operating other people's control planes actually teach, and which of those lessons transfer to a reader who is not SAP? The answer that the evidence keeps pointing at is a selection rule. Gardener's durable parts are the ones that have a one-to-one analogue in Kubernetes itself, and the parts it deleted are the ones that had no analogue or that fused two upstream concepts into one. The system's own vocabulary makes the rule visible: a tenant cluster is a pod, a host cluster is a node, the agent on the host cluster is a kubelet. Where that mapping held, the component survived seven years of releases. Where SAP invented something outside it, the code is gone, and the feature gate table gives the date.

Scope, and the evidence limit

This guide covers the architecture of the fleet control plane: where tenant control planes run, how they are exposed, scheduled, scaled, upgraded, isolated and recovered. It does not cover SAP's wider estate (HANA, ABAP, the Business Technology Platform application layer), Gardener's dashboard and CLI surface, or the product and commercial history. It also cannot quote SAP's own operational scale: this session's network policy resolved code hosts, package registries and one vendor price list, and refused every engineering blog, conference site, paper archive and status page, so there are no talks, papers or third-party measurements here, and no published incident review. Section 04 says what that absence does and does not permit a reader to conclude.

Figure 1 · The unit of scheduling is a control plane

places the control plane

tunnel opened from
the tenant side

Product team asks
for a cluster

Garden cluster
API, scheduler, controllers

Seed cluster
one per infrastructure and region
capacity: 200 tenants

Tenant control plane
apiserver plus two etcd,
pods in one namespace

Worker nodes in the
tenant's own account

places the control plane

tunnel opened from
the tenant side

Product team asks
for a cluster

Garden cluster
API, scheduler, controllers

Seed cluster
one per infrastructure and region
capacity: 200 tenants

Tenant control plane
apiserver plus two etcd,
pods in one namespace

Worker nodes in the
tenant's own account

A tenant asks for a cluster; what actually gets scheduled is a set of pods in a host cluster, with the tenant's worker nodes left in the tenant's own infrastructure account. Reconstructed from Gardener's architecture document and GEP-13.
Diagram source

Figure 2 · What got decided, and when

2017-2020Seeded fromKubernetessample-apiserver,SAP copyrightv1.0.0 ships withprovider code alreadyout of treeOne shared gatewayper host clusterreplaces per-tenantload balancers2021Adopted upstreamtunnel removed, ownreversed VPN insteadHost clusters getcapacity andallocatable, like anode2022-2023Certificate authorityrotation becomes anAPI operationFailure domainbecomes a per-tenantchoiceIn-house fusedautoscaler replacedby the upstream pair2024-2025Static cloudcredentials replacedby federated tokensThree advisories,tenant admin reachesthe host clusterIn-place node updatesfor bare metal andscarce GPUs2026Log collectionstandardised, then thestore swappedTenant clusters thathost their own controlplaneDated by the release that carried each decision
2017-2020Seeded fromKubernetessample-apiserver,SAP copyrightv1.0.0 ships withprovider code alreadyout of treeOne shared gatewayper host clusterreplaces per-tenantload balancers2021Adopted upstreamtunnel removed, ownreversed VPN insteadHost clusters getcapacity andallocatable, like anode2022-2023Certificate authorityrotation becomes anAPI operationFailure domainbecomes a per-tenantchoiceIn-house fusedautoscaler replacedby the upstream pair2024-2025Static cloudcredentials replacedby federated tokensThree advisories,tenant admin reachesthe host clusterIn-place node updatesfor bare metal andscarce GPUs2026Log collectionstandardised, then thestore swappedTenant clusters thathost their own controlplaneDated by the release that carried each decision
Each entry is dated by the first release that carried its decision record, or by the release that removed the mechanism, from the module proxy timestamps. The pattern worth noticing: the removals are later than the additions by three to five years.
Diagram source
02

How it is actually built

Kubernetes managing Kubernetes, with every operational concept renamed rather than reinvented. The naming is not a joke; it is the design method.

The core move is stated in the repository's own README: "The shoot clusters do not have dedicated master VMs. Instead, the control plane is deployed as a native Kubernetes workload into the seeds (the architecture is commonly referred to as kubeception or inception design). This does not only effectively reduce the total cost of ownership but also allows easier implementations for day-2 operations (like cluster updates or robustness) by relying on all the mature Kubernetes features and capabilities." Read that as a bet: if a tenant's control plane is an ordinary deployment, then every tool that already exists for deployments applies to it, including rolling updates, horizontal and vertical autoscaling, pod disruption budgets, topology spread and the scheduler.

The bet is carried through to the vocabulary. The README publishes the mapping directly, and it is worth reading as an architecture diagram in table form.

Kubernetes conceptGardener conceptWhat the analogue buys
PodShoot cluster (a tenant's cluster)Scheduling, eviction, capacity accounting
NodeSeed cluster (a host for control planes)Capacity and allocatable, drain, cordon
KubeletGardenlet, running in each seedPull-based reconciliation, no inbound path to the host
SchedulerGardener schedulerPlacement policy as data, not code
Controller manager, API serverGardener controller manager, extension API serverDeclarative spec and status for a cluster

Three consequences of that mapping are visible in the current code and documents, and they are the parts a reader should copy rather than the naming.

Availability is a scheduling property, not a replica count. The architecture document is explicit that control plane components "can be deployed with a replica count of 1 and only need to be scaled out when the control plane gets under pressure, but no longer for HA reasons", because the host cluster restarts them. That held until it did not: GEP-20, which first appears in the release of 11 August 2022, records that "many of the other critical control plane components including etcd are only offered with a single replica, making them susceptible to both node failure as well as zone failure causing downtimes", and introduces a per-tenant choice of failure domain. The design is still that redundancy is bought deliberately, per tenant, rather than applied uniformly.

The expensive shared resources are partitioned by name. Each tenant gets two etcd instances rather than one: the concepts document states the split exists so that "the critical etcd-main is not flooded by Kubernetes Events, as well as backup space is not occupied by non-critical data". At the documented default of 200 tenants per host cluster, that is 400 etcd instances in one Kubernetes cluster, each with a backup sidecar that ships snapshots to object storage and, per GEP-6, performs "data corruption checks ... prior to starting etcd". The network path in front of them is shared in the opposite direction: one load balancer per host cluster serves every tenant API server, using server name indication to route without terminating the tenant's TLS, with a deliberately stripped Istio installation as the gateway ("the default profile which is recommended for production deployment, is not suitable for the Gardener use case, as it offers more functionality than desired").

Infrastructure knowledge lives outside the core. GEP-1, the oldest proposal in the set and already implemented when version 1.0.0 shipped in February 2020, moved every cloud-specific and operating-system-specific behaviour into separate controllers, for a reason stated in one sentence: "Every change must be done centrally, requires to completely rebuild Gardener, and cannot be deployed individually. Similar to the motivation for Kubernetes to extract their cloud-specifics into dedicated cloud-controller-managers or to extract the container/storage/network/... specifics into CRI/CSI/CNI/..., we aim to do the same right now." The modern extension set covers five infrastructures that are continuously conformance-tested, and the extension contract is where the provider teams work without touching the core.

What left the platform is as instructive as what stayed. Version 1.0.0 in February 2020 shipped opinionated extras into every tenant cluster: an add-on chart bundle carrying a Kubernetes dashboard and an nginx ingress controller, plus a second bundle described in its own chart metadata as "temporarily experimental support for out-of-the-box installation of Kyma", another SAP project. Comparing the 2020 and 2026 release archives dates the retreat: the Kyma bundle was gone by release 1.9.0 on 25 August 2020, the bundled seed monitoring chart by 1.15.0 that January, after GEP-19 moved observability onto operators, and the last ingress add-on is being switched off behind feature gates named DisableNginxIngressInShoot, DisableNginxIngressInSeed and DisableNginxIngressInGarden, which appeared in release 1.142 in 2026. Six years to put down four add-ons: the platform kept the contracts and gave up the opinions.

Two divergence points matter when transferring this design. The first is that hosting control planes elsewhere is not always allowed: GEP-28, carried in the release of 13 November 2025, adds tenant clusters whose control plane runs on their own nodes, for environments "where the control plane must run side-by-side and cannot run in seed clusters", naming air-gapped operation, edge, firewalling and compliance. The proposal is careful to say it is not a kubeadm or k3s replacement, and it also fixes an awkwardness the team names itself: Gardener needed a conformant Kubernetes cluster to run in, "but so far there was no way of setting up or managing this initial cluster via Gardener itself", which the document calls "somewhat paradoxical". The second divergence is the recursion: host clusters can themselves be managed tenant clusters, which is how the fleet grows without a separate provisioning path.

Figure 3 · One host cluster, two hundred tenants

Seed cluster (one per infrastructure and region)

Namespace per tenant, x200

snapshots

scales down on
heartbeat loss

gardenlet
pulls desired state

Istio ingress gateway
one load balancer,
SNI routing, no TLS break

kube-apiserver

etcd-main

etcd-events

etcd-druid
operator

dependency-watchdog
meltdown protection

Object storage
per-tenant prefix

Seed cluster (one per infrastructure and region)

Namespace per tenant, x200

snapshots

scales down on
heartbeat loss

gardenlet
pulls desired state

Istio ingress gateway
one load balancer,
SNI routing, no TLS break

kube-apiserver

etcd-main

etcd-events

etcd-druid
operator

dependency-watchdog
meltdown protection

Object storage
per-tenant prefix

Everything inside the dotted boundary is ordinary Kubernetes workload, which is what makes day-two operations ordinary. Note the two shared singletons: the ingress gateway and the agent. Reconstructed from the etcd concept doc, the Istio operations doc and GEP-8.
Diagram source
03

The decisions that matter

Nine forks in the road, each taken in public, each with the condition that would make the rejected option right. Four are worth the long form.

Decision: where does a tenant's control plane physically run?

Chosen
  • As pods in a shared host cluster, no dedicated master machines
  • Stated reason: lower total cost of ownership, and day-two operations become ordinary Kubernetes operations
Rejected
  • Dedicated control plane machines per cluster, the model of most provisioning tools
  • Named in the README as "the main difference compared to many other OSS cluster provisioning tools"
Flips when
  • The control plane may not leave the tenant's network: air gap, edge, firewalled estate, or a compliance rule. SAP added exactly that flavour in GEP-28, eight years in, rather than bending the shared model

Decision: one load balancer per tenant, or one per host cluster?

Chosen
  • One load balancer for every tenant API server in a host cluster, routed by server name indication, TLS terminated by the tenant's own API server
  • Priced in the proposal: "ClassicLoadBalancer on AWS costs at minimum 17 USD / month"
Rejected
  • A load balancer per tenant, which the earlier design needed twice over, once for the API server and once for the VPN
  • Second stated reason is not cost but quota: "Quotas can limit the amount of LoadBalancers you can get per account / project, limiting the number of clusters you can host under a single account"
Flips when
  • A tenant needs its own address, its own quota or its own mitigation posture, or the infrastructure meters the shared gateway's traffic more expensively than its load balancers

Decision: keep the in-house autoscaler, or go back to upstream?

Chosen
  • Upstream horizontal and vertical autoscalers, used independently on the same workload (GEP-23, first carried 28 July 2023)
  • Stated benefit includes "reduces maintenance effort by reducing Gardener's custom code base"
Rejected
  • HVPA, SAP's own autoscaler that fused the two dimensions, alpha since release 0.31 and deleted in 1.109
  • Reason: it "poses severe algorithmic limitations which manifest as stability and efficiency issues in the field"
Flips when
  • The two upstream controllers genuinely fight over the same metric and you can prove it in production. Note that this took SAP roughly five years to conclude, and the workload in question is the single largest line in its compute bill

Decision: adopt the upstream fleet API, or keep your own?

Chosen
  • Gardener's own Shoot API, which fixes the make-up of the cluster: same version, same add-ons, same update and rotation behaviour on every infrastructure
  • Gardener contributed its machine abstraction upstream and was the first adopter of the Cluster API machine concepts
Rejected
  • Adopting Cluster API as the public interface
  • Reasons recorded in the concepts doc: provider-specific control plane resources mean "you cannot simply swap provider: foo with provider: bar", managed-service support is partly experimental, and the structure "doesn't always align naturally with how fully managed services architect their offerings"
Flips when
  • You need portability of the request rather than homogeneity of the result, or your platform is an aggregator of other people's managed services rather than an operator of clusters
DecisionChosenRejectedBecauseEvidence
Provider and operating system knowledgeOut-of-tree controllers behind a contractIn-tree branches per cloudCentral changes forced a full rebuild and could not be deployed individuallyGEP-1
Host-to-tenant connectivityA tunnel opened from the tenant side, built in-houseUpstream Konnectivity, adopted as alpha in release 1.6 and removed in 1.27The older design needed an extra load balancer per tenant and blocked private-only estatesGEP-14, gate table
etcd lifecycleA dedicated operator reconciling an Etcd resourceA mutating webhook that rewrote the StatefulSet to inject provider detailsThe webhook approach "restricts the operations on etcd, such as scale-up and upgrade"GEP-6
Placement of tenants onto hostsCapacity and allocatable on the host, enforced by the schedulerUnbounded placement"Seeds have a practical limit of how many shoots they can accommodate"GEP-13
Who decides when a version diesClassifications plus a machine-readable expiry date per versionManual promotion through stagesUpstream ships a minor "roughly every three months" and maintains three; manual movement "is cumbersome"GEP-5, GEP-32
Credentials used to call the infrastructureShort-lived federated tokens issued by the platformLong-lived service keys supplied by the tenantSuch credentials are "often non-expiring", widely copied, and their accidental invalidation stops reconciliationGEP-26
Changing a worker node's configurationIn-place update as an alternative strategyRolling replacement onlyBare metal has long boot and local disks; and virtual machine types can be "scarce ... because of an ongoing capacity crunch", naming GPUsGEP-31
Log storage for every control planeOpenTelemetry collection into VictoriaLogsStaying on a fork of the previous engineThe fork "maintains only security updates, thus leading to no new features or improvements getting integrated"GEP-35
Recovering a tenant whose host cluster is unreachableOwnership passed by a DNS text record, checked by every participantSynchronisation objects written to the backup bucketWith owner records "such sync objects are no longer needed", and split brain is prevented without the source host participatingGEP-17

Read the table column by column rather than row by row and the selection rule appears. Every chosen option on the left has a Kubernetes analogue: a contract with out-of-tree controllers is the cloud controller manager, capacity and allocatable is the node status, an operator reconciling a custom resource is the operator pattern, independent horizontal and vertical autoscalers are upstream components. Every rejected option is either an invention with no analogue (a fused two-dimensional autoscaler, synchronisation files in a bucket, a webhook that rewrites another controller's output) or an upstream component used against the grain. The two mechanisms Gardener took out of its own feature gate table without ever promoting them to general availability, Konnectivity and HVPA, are one of each kind.

Figure 4 · Which shape of fleet control plane

No

Yes

No

Yes

Tens

Hundreds

More than one
infrastructure, and
clusters must match?

Buy the managed service
and pay per cluster hour

May the control plane
leave the tenant's
network?

Control plane on the
tenant's own nodes,
managed remotely

Tens of clusters,
or hundreds?

Dedicated control plane
machines per cluster

Control planes as pods
in shared host clusters

No

Yes

No

Yes

Tens

Hundreds

More than one
infrastructure, and
clusters must match?

Buy the managed service
and pay per cluster hour

May the control plane
leave the tenant's
network?

Control plane on the
tenant's own nodes,
managed remotely

Tens of clusters,
or hundreds?

Dedicated control plane
machines per cluster

Control planes as pods
in shared host clusters

The three leaves are real products, and the questions are the ones Gardener's own proposals answer in order. Derived from GEP-28, the Cluster API comparison and GKE's price list.
Diagram source
04

What broke, and what the record will not tell you

No public incident review exists for this fleet. What does exist is a vulnerability record, a set of design documents written against named failure modes, and a component whose entire purpose is to stop a recovery mechanism from making things worse.

Read this before the cards

SAP publishes no postmortems for the landscapes it runs on Gardener, and this session could not reach any engineering blog that might carry one. Three of the four cards below are built from vendor security advisories, which are incident records with a known root cause and a known blast radius but no timeline and no detection story. The fourth is built from design documents that state the failure they exist to prevent. That means a reader can trust the mechanisms and the design rules here, and should treat every claim about frequency, duration or detection time as unavailable rather than as reassuring.

Advisory

The tenancy boundary is a validation function

AssumptionA tenant with administrative rights over their own project cannot reach the cluster that hosts their control plane.
What happenedTwice in May 2025, once through secret validation that could be bypassed and once through metadata injection on a project secret, "a user with administrative privileges for a Gardener project" could "obtain control over the seed cluster(s) where their shoot clusters are managed". The CVSS vectors are identical and both carry a changed scope.
Blast radiusThe host cluster, and therefore every other tenant's control plane on it. One advisory states it "affects all Gardener installations no matter of the public cloud provider(s) used". Fixes went out on four maintenance lines at once, 1.116.4, 1.117.5, 1.118.2 and 1.119.0.
FixInput validation on the path into the host cluster, shipped as patch releases rather than as an architectural change.
Design ruleWhen tenants share a host, the isolation boundary is whatever code validates tenant input before it is rendered into host-side objects. Budget for that code as a security boundary: review it like a parser, fuzz it, and keep enough supported release lines alive to patch it everywhere at once.
SourceGHSA-3hw7-qj9h-r835, GHSA-9x73-87fh-54w9, published 19 May 2025
Advisory

Tenant configuration reached a template engine

AssumptionProvider extensions can pass a tenant's providerConfig into the infrastructure tooling they drive.
What happenedIn September 2025 the same escalation arrived through a different door: where Terraform is used for infrastructure provisioning, code injection was possible in the AWS, Azure, GCP and OpenStack extensions, again ending in control over the host cluster.
Blast radiusFour provider extensions, each with its own fixed version (1.64.0, 1.55.0, 1.46.0, 1.49.0). The out-of-tree design that makes providers independently deployable also means a class of defect has to be fixed four times.
FixSanitising tenant-supplied configuration before it reaches the provisioning tool.
Design ruleAny tenant field that ends up inside a generated program, template or shell context is an injection site, and extracting components into separate repositories multiplies the number of places that site exists. Put the validation in the shared contract, not in each extension.
SourceGHSA-227x-7mh8-3cf6, published 25 September 2025
Design record

The recovery mechanism eats healthy machines

AssumptionA node that stops renewing its lease is unhealthy and should be replaced.
What happenedThe scenario the dependency watchdog documents: in a cluster with several hundred nodes, a broken network path (the document's example is a NAT gateway) stops every kubelet from reaching its API server. The controller manager moves the nodes to an unknown state, and the machine controller manager, doing its job, "will begin to replace the unhealthy machine(s) with new ones", which "results in undesired downtimes for customer workloads that were running on these otherwise healthy nodes".
Blast radiusEvery node of the affected tenant, with the damage caused by the platform rather than by the original fault. No duration is published.
FixA separate component probes the API server and then counts expired node leases; past a configured fraction it scales the dependent controllers down and annotates them, in the code's own words, with meltdown protection, then scales them back up when leases recover.
Design ruleAny controller that acts on the absence of a signal needs a second opinion about whether the signal path is healthy. Two probes beat one threshold: ask "can I reach the server" and "can the fleet reach the server", and treat disagreement as a reason to stop acting, not a reason to act faster.
Sourcedependency-watchdog prober concept, checked 1 October 2026
Design record

The upstream you forked stops moving

AssumptionForking an open-source log store at its last permissively licensed commit buys time at no cost.
What happenedAfter the log engine relicensed, Gardener's observability stack moved to a fork of the last Apache-2.0 version. GEP-35 records the consequence: the fork "maintains only security updates, thus leading to no new features or improvements getting integrated", which left the fleet unable to adopt the collection protocol it now wants.
Blast radiusEvery tenant control plane with logging enabled, plus the host and garden clusters, and a migration that has to move accumulated data without losing it.
FixCollection moved to OpenTelemetry first (GEP-34), which made the storage engine replaceable, and then storage moved to VictoriaLogs. The collector's gate reached beta in release 1.136 and the new log store's gate opened in 1.137, twelve days later in February 2026.
Design ruleA fork of a dependency is a dated loan, not a decision. Pay it down by putting a standard interface between you and the dependency before the fork goes stale; the interface, not the fork, is what makes the next swap cheap.
SourceGEP-35, carried from release 1.132.0, 13 November 2025

The four cards group into two failure classes, and both are properties of the architecture rather than bugs in it. The first class is shared-host escalation: packing two hundred tenants' control planes into one cluster converts an input-validation defect into a cross-tenant compromise, which is why all three 2025 advisories end with the same clause about obtaining control over the host. The second class is a controller acting on stale or missing information, which appears as machines being replaced during a network fault, as split brain during a control plane migration, and in GEP-20 as a single etcd replica deciding a whole tenant's availability. Gardener's answers to the second class are consistent and worth stealing: a second probe before acting, an ownership record that no participant can forge, and an explicit per-tenant choice of failure domain.

Figure 5 · The meltdown path, and where it is cut

dependencywatchdogmachine controllercontroller managerkube-apiserver(seed)Kubelets (shoot)dependencywatchdogmachine controllercontroller managerkube-apiserver(seed)Kubelets (shoot)healthy machinessurvive the faultlease renewals stop (networkfault)leases expirednodes marked Unknownnode status changedgrace period, thenreplace machinesprobe 1, can I reach the apiserver? yesprobe 2, expired leases above threshold? yesscale to zero, annotate meltdownprotectionleases recoverscale back up, removeannotation
dependencywatchdogmachine controllercontroller managerkube-apiserver(seed)Kubelets (shoot)dependencywatchdogmachine controllercontroller managerkube-apiserver(seed)Kubelets (shoot)healthy machinessurvive the faultlease renewals stop (networkfault)leases expirednodes marked Unknownnode status changedgrace period, thenreplace machinesprobe 1, can I reach the apiserver? yesprobe 2, expired leases above threshold? yesscale to zero, annotate meltdownprotectionleases recoverscale back up, removeannotation
The platform's own recovery logic is the dangerous actor once the heartbeat path breaks, so the watchdog disables it rather than the nodes. Reconstructed from the prober concept document.
Diagram source
05

Numbers you can plan against

Everything quantitative in the corpus, with the date attached and the derivation shown where the number is mine rather than theirs.

MetricValueKindContextAs ofSource
Tenant clusters per host cluster200Measured in configDefault capacity.shoots in the shipped agent configuration; the scheduler refuses to place more2026-09gardenlet config
etcd instances per host cluster at default capacity400DerivedTwo per tenant, main and events, times 200 tenants; each with a backup sidecar2026-09etcd concepts
Load balancer cost avoided by sharing one gateway~$3,380/moDerived199 tenants times the proposal's figure of 17 USD per month for a classic load balancer; the figure predates 2020 and current prices differ2019GEP-8
Managed-service price for one control plane$0.10/hrVendor list priceGKE's flat cluster management fee, about $73 per 730-hour month, independent of cluster size or topology2026-10GKE pricing
Penalty for staying on an unsupported version+$0.50/hrVendor list priceGKE's extended support fee, taking the total to $0.60 per cluster per hour; Gardener instead expires the version2026-10GKE pricing
Minor releases of the control plane152MeasuredFrom 1.0.0 on 6 February 2020 to 1.152.0 on 24 September 2026, about one every 16 days, against a stated target of "roughly every other week"2026-09module proxy, process doc
Released versions in total759MeasuredIncluding patches, and including the 0.x line back to 29 November 20192026-10module proxy
Release lines kept patchable3Stated policy"Hotfixes are usually maintained for the latest three minor releases"; the May 2025 advisories were fixed on four lines2026-09process doc
Feature gates documented69DerivedDistinct feature names in the gate table, counted by parsing it; 39 carry a general-availability row2026-09gate table
Median time from alpha to general availability15 releasesDerivedAbout 225 days, over the features with both an alpha and a general-availability row; slowest was 57 releases and 904 days2026-09gate table, release dates
Upstream Kubernetes cadence the fleet absorbs3 minors/yrStated"The Kubernetes community releases minor versions roughly every three months and usually maintains three minor versions"; control plane and workers are kept on the same version2020GEP-5
Infrastructures under continuous conformance test5StatedAWS, Azure, GCP, OpenStack and Alibaba Cloud, certified up to Kubernetes 1.352026-09README
Variation in control plane resource needs100xStated"Compute resources required by ShootKapis of different shoots vary by two orders of magnitude", which is why one fixed size per tenant is not an option2023GEP-23
What to plan against

Two of these numbers do the most work in a design review. The first is the packing ratio: a host cluster configured for 200 tenant control planes is the difference between a fleet whose cost scales with clusters and one whose cost scales with load, and the managed-service price list is the alternative you are being measured against. The second is the cadence pair, three upstream minor versions a year arriving into a platform that ships its own minor every sixteen days and keeps three lines patchable; that ratio, not the cluster count, is what sets the size of the team. Everything in the cost column is a list price or a figure inside a proposal, not an audited bill, and the oldest of them is seven years old.

06

The evidence wall

Every source behind this page, graded and dated. The shape of this wall is itself a finding: nineteen decision records, eleven repository artefacts, three advisories, one price list, and no blog post, talk or paper, because the network policy in this session resolved code hosts and registries only.

Advisory Gardener2025-05-19

GHSA-3hw7-qj9h-r835: bypassing project secret validation

A tenant project administrator could obtain control over the host clusters running their control planes, in every installation regardless of cloud provider. Fixed on four maintenance lines simultaneously.

Carry forwardIn a shared-host fleet, the tenancy boundary is the validation code on the path from tenant input to host-side objects.
GitHub Advisory Database record
Advisory Gardener2025-05-19

GHSA-9x73-87fh-54w9: metadata injection on a project secret

The same escalation through a second path, this time in the host-cluster agent and scoped to installations using the GCP provider extension. Identical severity vector, with a changed scope.

Carry forwardTwo independent paths to the same escalation in one month is a signal about the class of defect, not about the two bugs.
GitHub Advisory Database record
Advisory Gardener2025-09-25

GHSA-227x-7mh8-3cf6: code injection through provider configuration

Where Terraform drives infrastructure provisioning, tenant-supplied configuration could be injected, again ending in control over the host cluster. Four provider extensions each needed their own fixed version.

Carry forwardOut-of-tree extensions multiply the places a shared defect class must be fixed; put the sanitisation in the contract.
GitHub Advisory Database record
Decision record Gardenerin v1.0.0, 2020-02

GEP-1: extensibility and extraction of cloud-specific knowledge

The founding decision: provider and operating-system behaviour leaves the core for separate controllers behind a contract, explicitly modelled on Kubernetes extracting its own cloud, container, storage and network specifics.

Carry forwardThe trigger to extract is not code size; it is that a provider change forces a rebuild of the whole platform.
GEP-1
Decision record Gardenerin v1.0.0, 2020-02

GEP-8: one load balancer for every tenant API server

Routes every tenant API server in a host cluster through one gateway using server name indication, with the tenant's API server still terminating TLS. Prices the rejected option and names the quota limit behind it.

Carry forwardPer-tenant infrastructure objects are a quota problem before they are a cost problem, and quotas are harder to raise than budgets.
GEP-8
Decision record Gardenerin v1.4.0, 2020-05

GEP-11: invert host-to-tenant connectivity with the upstream proxy

Documents the pre-2020 design, a tunnel needing a load balancer in every tenant cluster, and proposes adopting the upstream proxy instead. The gate it added was removed sixteen months later.

Carry forwardRead an adopted upstream component's exit date, not only its entry date, before copying the decision.
GEP-11
Decision record Gardenerin v1.21.0, 2021-04

GEP-14: the tunnel is opened from the tenant side

Supersedes both the old tunnel and the adopted upstream proxy, and states the intention to delete the upstream code once validated. The non-goals are unusually honest: not a low latency path, no availability promise.

Carry forwardWriting down what a connectivity path is explicitly not for is what stops the next team from loading it with traffic it cannot carry.
GEP-14
Decision record Gardenerin v1.14.0, 2020-12

GEP-13: host clusters get capacity and allocatable

Makes placing tenants onto hosts a scheduling problem using the fields a node publishes, because "seeds have a practical limit of how many shoots they can accommodate".

Carry forwardGive your host a declared capacity early; an unbounded packing ratio is discovered during an incident.
GEP-13
Decision record Gardenerin v1.0.0, 2020-02

GEP-6: etcd gets an operator instead of a webhook

Replaces a mutating webhook that rewrote the etcd StatefulSet to inject provider details, because that "restricts the operations on etcd, such as scale-up and upgrade".

Carry forwardA webhook that edits another controller's output is a lifecycle liability; model the thing you are managing as its own resource.
GEP-6
Decision record Gardenerin v1.41.0, 2022-02

GEP-18: rotating a tenant's certificate authority as an API call

Turns the most dangerous manual operation in a cluster into a staged, declarative one, which is the only form that survives being multiplied by a fleet.

Carry forwardAt fleet scale, a credential rotation runbook is not an operation; it has to become an API with states.
GEP-18
Decision record Gardenerin v1.53.0, 2022-08

GEP-20: failure-domain tolerance as a per-tenant choice

States plainly that critical components including etcd ran with a single replica, "making them susceptible to both node failure as well as zone failure", and makes isolation a per-tenant selection.

Carry forwardUniform redundancy is the expensive default; make the failure domain a field on the tenant's resource.
GEP-20
Decision record Gardenerin v1.76.0, 2023-07

GEP-23: delete the in-house autoscaler, use the upstream pair

The clearest cost document in the set: tenant API server compute is "a major part of Gardener's overall compute cost", requirements vary by two orders of magnitude, and the home-grown fused autoscaler had "severe algorithmic limitations".

Carry forwardFusing two upstream controllers into one is the invention most likely to be deleted; prefer composing them and measuring the fight.
GEP-23
Decision record Gardenerin v1.93.0, 2024-04

GEP-26: federated tokens instead of tenant cloud keys

Replaces long-lived infrastructure credentials with short-lived tokens the cloud provider trusts through federation, naming the operational failure as well as the security one.

Carry forwardCredential expiry is an availability problem for a reconciler, which is the argument that gets rotation funded.
GEP-26
Decision record Gardenerin v1.112.0, 2025-02

GEP-31: update nodes in place instead of replacing them

Rolling replacement assumes interchangeable machines; the proposal lists what breaks that, including long bare-metal boot, local disks, and machine types scarce during a capacity crunch, with GPUs named.

Carry forwardHardware scarcity is an architectural input: when you cannot get the replacement, immutable infrastructure stops being a free choice.
GEP-31
Decision record Gardenerin v1.132.0, 2025-11

GEP-28: tenant clusters that host their own control plane

Adds the model the architecture was built to avoid, for air-gapped, edge and compliance cases, and so that Gardener can bootstrap the cluster it runs in, which the document calls "somewhat paradoxical" to have needed third-party tools for.

Carry forwardA platform that cannot create its own first instance has a bootstrap dependency it will eventually pay to remove.
GEP-28
Decision record Gardenerin v1.132.0, 2025-11

GEP-35: replacing a forked log store

Records why the fleet sat on a fork of a relicensed log engine, what that cost, and why collection was standardised first so the storage engine became swappable.

Carry forwardStandardise the interface before the fork goes stale; the interface is what makes the second migration cheap.
GEP-35
Decision record Gardenerin v1.41.0, 2022-02

GEP-17: recovering a tenant whose host cluster is unreachable

Ownership passes through a DNS text record that every participant checks, including the etcd backup sidecar, so the source host need not cooperate for split brain to be avoided. An earlier design using files in the backup bucket is recorded as no longer needed.

Carry forwardFor a stateful tenant, the recovery primitive is a forgery-resistant ownership record plus migration, not repair in place.
GEP-17
Decision record Gardenerin v1.0.0, 2020-02

GEP-5: a version lifecycle with four classifications

Sets the policy for marking versions preview, supported, deprecated or expired, against an upstream that ships a minor roughly quarterly and maintains three.

Carry forwardPublish the lifecycle as data your tenants can query, or you will negotiate every upgrade individually.
GEP-5
Decision record Gardenerin v1.137 era, 2026

GEP-32: the whole version lifecycle, scheduled in advance

Extends classifications with scheduled transitions, because moving versions through stages by hand "is cumbersome".

Carry forwardAn expiry date is the cheapest forcing function a platform has; make it declarative and let the calendar do the arguing.
GEP-32
Decision record Gardenerchecked 2026-10-01

Why not the upstream Cluster API

A standing comparison: Cluster API harmonises getting to clusters, Gardener harmonises the clusters themselves. Lists the blockers, including provider-specific control plane resources and experimental managed-service providers.

Carry forwardDecide whether you are standardising the request or the result; the two need different APIs and only one gives you homogeneity.
Cluster API relation doc
Decision record Gardenerchecked 2026-10-01

The enhancement process moved out of the code repository

Project-wide proposals now live in a dedicated repository under a technical steering committee, modelled on the upstream Kubernetes process. The pointer files left in the code repository date the move.

Carry forwardWhen decisions outlive the release they shipped in, move them out of the release artefact and give them their own lifecycle.
gardener/enhancements README
Source Gardenerv1.152.0, 2026-09

README: kubeception, and the concept mapping

States the central decision (no dedicated master machines, control planes as workload, lower total cost of ownership) and publishes the concept mapping the rest of the architecture follows.

Carry forwardNaming your fleet concepts after the system you already operate is a design constraint, not documentation.
README at v1.152.0
Source Gardenerv1.152.0, 2026-09

Architecture concept: one host cluster per infrastructure and region

The sizing statement ("hundreds or even thousands of clusters"), the placement rule, and the claim that single replicas are acceptable because the host cluster watches them.

Carry forwardIf the host restarts your component, availability becomes a scheduling property; write that assumption down so you notice when it stops holding.
architecture.md
Source Gardenerv1.152.0, 2026-09

The feature gate table, which dates every mechanism

Sixty-nine documented features with the release each entered and left, including the adopted upstream tunnel (alpha in 1.6, removed in 1.27) and the in-house fused autoscaler (alpha since 0.31, removed in 1.109).

Carry forwardA dated gate table is an audit trail for architecture; keep one and your deprecations stop being folklore.
feature_gates.md
Source Gardenerv1.152.0, 2026-09

Shipped agent configuration: capacity of 200 tenants

The default packing ratio is a line of YAML in the example component configuration, and the scheduler treats it as the Kubernetes scheduler treats a node's capacity.

Carry forwardThe most important number in a multi-tenant platform is usually a default in a config file; find it before you model cost.
20-componentconfig-gardenlet.yaml
Source Gardenerv1.152.0, 2026-09

etcd concept: two stores per tenant, and a corruption check on start

Events are split away from the critical store so neither the hot path nor the backup bucket is flooded by low-value data; a sidecar handles snapshots, defragmentation and validation.

Carry forwardSplit a shared datastore by the value of the data, not only by size; the cheap half is what fills your backups.
etcd.md
Source Gardenerv1.152.0, 2026-09

Istio, installed for exactly one job

Lists everything deliberately not deployed from the upstream default profile: telemetry, the egress gateway, the sidecar injector, the addons.

Carry forwardAdopting a large dependency is survivable if you write down the subset you use and refuse the rest in configuration.
istio.md
Source Gardenerv1.152.0, 2026-09

Release process: every other week, three lines patchable, a named rota

Carries the cadence target, the hotfix window, and a rota naming a responsible engineer for each version through late 2026. Fourteen names rotate.

Carry forwardA fleet platform's release cadence has to be a rota with names in it, because the upstream it tracks will not slow down.
process.md
Source Gardenerchecked 2026-10-01

dependency-watchdog: the prober and the weeder

A component built to stop cascading failure by "conservatively scaling down dependent configured resources", documenting the scenario, the two probes, the threshold and the meltdown-protection annotation.

Carry forwardShip the brake with the engine: automated recovery needs a documented, annotated way to be switched off while a fault lasts.
prober.md
Source Gardenerv1.152.0, 2026-09

The notice file that dates the origin

"Copyright 2017-2019 SAP SE or an SAP affiliate company" still ships in the current release, alongside the record that the project was seeded from Kubernetes' sample API server.

Carry forwardCopyright and notice files are the cheapest provenance evidence in a repository, and they outlive the blog posts.
NOTICE.md
Source Go module proxychecked 2026-10-01

759 released versions, each with a timestamp

The proxy lists every published version with its tag time, which dates this guide's timeline, and serves the full source archive of any version, which is how the 2020 and 2026 trees were compared.

Carry forwardFor any Go project, the module proxy is a dated archive of every release; it answers "when did this appear" without repository access.
proxy.golang.org version list
Source pkg.go.devchecked 2026-10-01

Release dates, per year

The versions tab dates every release: 49 in 2020, then 70, 77, 79, 75, 92 and 79 so far in 2026.

Carry forwardRelease cadence over years is the most honest public signal of whether a platform is staffed.
Version history
Decision record Gardenerin v1.132.0, 2025-11

GEP-34: standardise observability collection before swapping the store

Puts an OpenTelemetry collector in every tenant control plane because the stack "still relies on vendor specific format and protocols", and refuses a flag day: the migration is phased and the old components are not decommissioned immediately.

Carry forwardReplace the protocol first and the storage second; doing it in that order turns the next migration into a configuration change.
GEP-34
Vendor Google Cloudchecked 2026-10-01

What a managed control plane costs per hour

"A flat cluster management fee of $0.10 per cluster per hour ... applies to all GKE clusters irrespective of the mode of operation, cluster size, or topology", plus $0.50 per hour more for clusters left on a version in extended support.

Carry forwardThe build-or-buy line for a fleet platform is a price per cluster hour; and note that the market charges for falling behind on versions.
GKE pricing
07

Build a miniature, then productionise it

Six rungs. The first three fit in an evening on a laptop; the line from toy to production-shaped is crossed at rung four, where failure handling starts.

Host one control plane as a workload

In a local cluster, run a second Kubernetes API server and its own etcd as ordinary pods in a namespace, with their own certificate authority, and point a kubeconfig at it through a service.

Done when: kubectl --kubeconfig tenant.yaml get ns answers from the pod-hosted API server.  Teaches: why the inception model makes day-two operations ordinary.

Attach a worker that lives somewhere else

Join one node to that control plane from outside the host cluster, with the tunnel opened from the worker side so no inbound path to the worker network is required.

Done when: the node is Ready and a pod on it is reachable, with no inbound firewall rule on the worker side.  Teaches: the connectivity inversion in GEP-14, and why it is explicitly not a high throughput path.

Put twenty tenants behind one address

Add a gateway that routes to the right tenant API server by server name indication, without terminating TLS, and give each tenant its own hostname.

Done when: twenty kubeconfigs with twenty hostnames and one IP address all authenticate with client certificates.  Teaches: the load balancer and quota arithmetic from GEP-8.

Declare capacity and let a scheduler place tenants

Model the host's capacity and allocatable tenant count, and write a placement loop that refuses to exceed it and spreads across hosts. Then fill a host and watch the next request stay pending.

Done when: the twenty-first tenant stays pending with a reason, rather than degrading the twenty already there.  Teaches: that a packing ratio is a scheduling constraint, which is the lesson of GEP-13.

Break the heartbeat path and survive your own recovery

Cut connectivity between the workers and their control plane, and watch what your node lifecycle controller does. Then add the second probe: count expired leases, and scale the replacing controller to zero past a threshold.

Done when: a thirty-minute network fault destroys no healthy machines, and the controllers come back by themselves afterwards.  Teaches: the meltdown-protection pattern, and that absence of signal is not evidence of failure.

Move a tenant to another host while pretending the first is gone

Back the tenant's etcd to object storage, then restore it on a second host without the first host cooperating. Add an ownership record that both sides check before writing, and test the case where the old host comes back.

Done when: the tenant works on the new host, and the revived old host refuses to reconcile because it lost the ownership check.  Teaches: why GEP-17 chose a record no participant can forge over files in a shared bucket.

08

Keep hunting

This page was assembled without access to a single engineering blog. These are the techniques that replaced them, and they work on any company that ships Go.

Dated archaeology without repository access

  • curl -sL https://proxy.golang.org/<module>/@v/list
  • curl -s https://proxy.golang.org/<module>/@v/v1.0.0.info
  • curl -sL https://proxy.golang.org/<module>/@v/v1.0.0.zip -o old.zip
  • https://pkg.go.dev/<module>?tab=versions

Finding when a decision entered or left

  • raw.githubusercontent.com/<org>/<repo>/<tag>/<path>
  • bisect the tag list for the first release containing a design doc
  • grep the feature gate or deprecation table for Since and Until columns
  • diff the docs tree of the oldest and newest release archive

Incident evidence when there is no postmortem

  • raw.githubusercontent.com/github/advisory-database/main/advisories/github-reviewed/<year>/<month>/<GHSA>/<GHSA>.json
  • "obtain control over" OR "privilege escalation" <product> advisory
  • pkg.go.dev/vuln/ for the Go vulnerability database entry

The decision layer, in its own repository

  • <project> enhancement proposal OR GEP OR KEP "alternatives considered"
  • docs/proposals OR geps/ OR docs/adr in the project repository
  • "non-goals" <project> design proposal
09

References

  1. Gardener, README at v1.152.0 gardener/gardener. Checked 2026-10-01.
  2. Gardener, Architecture concept gardener/gardener, v1.152.0. Checked 2026-10-01.
  3. Gardener, NOTICE gardener/gardener, v1.152.0, copyright 2017-2019 SAP SE. Checked 2026-10-01.
  4. Gardener, Enhancements repository README gardener/enhancements. Checked 2026-10-01.
  5. GEP-1, Gardener extensibility and extraction of cloud-specific knowledge In releases from v1.0.0, 2020-02-06. Checked 2026-10-01.
  6. GEP-5, Gardener versioning policy In releases from v1.0.0. Checked 2026-10-01.
  7. GEP-6, Integrating etcd-druid with Gardener In releases from v1.0.0. Checked 2026-10-01.
  8. GEP-8, Shoot API server via SNI In releases from v1.0.0. Checked 2026-10-01.
  9. GEP-11, Utilize API server network proxy to invert seed-to-shoot connectivity In releases from v1.4.0, 2020-05-07. Checked 2026-10-01.
  10. GEP-13, Automated seed management In releases from v1.14.0, 2020-12-09. Checked 2026-10-01.
  11. GEP-14, Reversed cluster VPN In releases from v1.21.0, 2021-04-22. Checked 2026-10-01.
  12. GEP-17, Shoot control plane migration bad case scenario In releases from v1.41.0, 2022-02-25. Checked 2026-10-01.
  13. GEP-18, Shoot certificate authority rotation In releases from v1.41.0. Checked 2026-10-01.
  14. GEP-20, Highly available shoot control planes In releases from v1.53.0, 2022-08-11. Checked 2026-10-01.
  15. GEP-23, Autoscaling the shoot API server with independent HPA and VPA In releases from v1.76.0, 2023-07-28. Checked 2026-10-01.
  16. GEP-26, Workload identity In releases from v1.93.0, 2024-04-19. Checked 2026-10-01.
  17. GEP-28, Self-hosted shoot clusters In releases from v1.132.0, 2025-11-13. Checked 2026-10-01.
  18. GEP-31, In-place node updates In releases from v1.112.0, 2025-02-07. Checked 2026-10-01.
  19. GEP-32, Version classification lifecycle gardener/enhancements. Checked 2026-10-01.
  20. GEP-34, Observability 2.0, OpenTelemetry operator and collectors In releases from v1.132.0. Checked 2026-10-01.
  21. GEP-35, Observability 2.0, VictoriaLogs In releases from v1.132.0. Checked 2026-10-01.
  22. Gardener, Relation between the Gardener API and Cluster API gardener/gardener, v1.152.0. Checked 2026-10-01.
  23. Gardener, Feature gates gardener/gardener, v1.152.0. Checked 2026-10-01.
  24. Gardener, Example gardenlet component configuration gardener/gardener, v1.152.0. Checked 2026-10-01.
  25. Gardener, etcd concept gardener/gardener, v1.152.0. Checked 2026-10-01.
  26. Gardener, Istio operations gardener/gardener, v1.152.0. Checked 2026-10-01.
  27. Gardener, Releases, features, hotfixes gardener/gardener, v1.152.0. Checked 2026-10-01.
  28. Gardener, Dependency watchdog prober concept gardener/dependency-watchdog. Checked 2026-10-01.
  29. Gardener, Dependency watchdog README gardener/dependency-watchdog. Checked 2026-10-01.
  30. GHSA-3hw7-qj9h-r835 (CVE-2025-47283), bypassing project secret validation GitHub Advisory Database, published 2025-05-19. Checked 2026-10-01.
  31. GHSA-9x73-87fh-54w9 (CVE-2025-47284), metadata injection for a project secret GitHub Advisory Database, published 2025-05-19. Checked 2026-10-01.
  32. GHSA-227x-7mh8-3cf6 (CVE-2025-59823), code injection through provider configuration GitHub Advisory Database, published 2025-09-25. Checked 2026-10-01.
  33. Go module proxy, version list for github.com/gardener/gardener Checked 2026-10-01.
  34. pkg.go.dev, version history for github.com/gardener/gardener Checked 2026-10-01.
  35. Google Cloud, Google Kubernetes Engine pricing Checked 2026-10-01.