concept

Noisy Neighbour

also called Resource Contention, Co-Tenant Interference, Shared-Fate Degradation

One workload degrading others by consuming a shared resource - and specifically the resources nobody set limits on, such as control-plane capacity, node network bandwidth and local disk.

kubernetesmulti-tenancyquotasthrottlingisolation

Shared infrastructure achieves utilisation by placing multiple workloads on the same hardware. Isolation is provided for the resources someone thought to limit, and not for the others — and the others are where the incidents come from.

CPU and memory are limited almost everywhere. Node network bandwidth, local disk throughput, file descriptors, connection pool capacity in shared services, and the platform's own control-plane API are frequently unlimited, and a single workload can exhaust any of them while remaining entirely inside its declared quota.

Why it matters

The victim has no visibility into the cause. A service experiencing elevated latency because a co-tenant is saturating node disk sees only its own symptoms, and the investigation proceeds through its own code, its own dependencies and its own configuration before anyone considers the neighbour.

The diagnostic difficulty is the real cost. A workload that fails cleanly is a small incident; one that degrades an unrelated team's service through an unmonitored shared resource can consume days, and the correlation is only visible from the platform's perspective.

Implementation patterns

  • Resource requests and limits on every workload, enforced by admission policy — a workload with no request is scheduled as though it needs nothing, which is the root of most of these incidents.
  • Understand the asymmetry between CPU and memory limits. CPU limits throttle; memory limits kill. CPU throttling produces latency spikes invisible in average utilisation and is a common cause of unexplained tail latency, and setting the CPU limit equal to the request is a frequent, damaging default that prevents a workload from using idle capacity.
  • Namespace or team quotas capping total consumption.
  • Priority classes with preemption, so critical work displaces batch work under pressure.
  • Node pools by workload class, separating latency-sensitive serving from batch and from anything with unpredictable resource behaviour.
  • Per-tenant control-plane rate limits, since a team's controller in a hot loop can degrade the API server for everyone — a shared resource that is routinely overlooked.
  • Reservations and a dedicated pool for the platform's own components. Ingress controllers, log shippers, metrics agents and storage drivers run everywhere, and a tenant that starves them removes the platform's ability to observe and route while every tenant quota is respected.
  • Per-tenant attribution in observability, so the platform can identify the cause even when the victim cannot.

Industry example

The pattern is universal in shared Kubernetes platforms and is the principal argument for node pools and, in stronger cases, cluster-per-tenant models. The recurring published finding from platform teams is consistent: the isolation gaps that cause incidents are not CPU and memory — those are limited — but the resources with no natural limit, particularly node-local disk and network, and the platform's own control plane.

The same reasoning drives shuffle sharding and cell-based architecture in multi-tenant services: where tenant behaviour cannot be constrained, the structural response is to bound how many others any one tenant can affect.

Failure scenarios

  • Workloads with no resource requests, scheduled as though free.
  • CPU limits equal to requests, throttling workloads that could have used idle capacity and producing mysterious tail latency.
  • Unlimited node-local disk or network, where a log-heavy or data-heavy workload degrades its neighbours.
  • An unthrottled controller hammering the platform API and degrading it for every tenant.
  • Platform components sharing nodes with tenant workloads, without reservations.
  • A single tenant exhausting a shared service's connection pool — a database, a cache, a message broker.
  • No per-tenant attribution, so the platform team cannot identify the cause any faster than the victim can.
  • Quotas set once and never revisited as workloads grow.

Trade-offs

Strong isolation costs utilisation, which is the entire reason for sharing infrastructure in the first place. Dedicated node pools fragment capacity; generous limits leave headroom unused; cluster-per-tenant multiplies operational burden and makes the upgrade problem substantial.

Tight limits also cause their own incidents: a workload throttled or killed by a limit set from an out-of-date estimate fails for reasons unrelated to its own behaviour, and the failure is attributed to the platform.

The trade is utilisation and operational simplicity against the blast radius of one workload's behaviour. The proportionate answer depends on the threat model: internal teams that trust one another need protection from accidents rather than attacks, which is far cheaper — and where hostile or compliance-separated tenants are involved, the cheap option is not available and the utilisation loss is simply the price.

Interview question

"Team A's service gets slow every night at 2am and their code has not changed. Tell me how you would investigate, what you would look at that they cannot see, and what you would put in place afterwards so that the next occurrence is diagnosed in ten minutes rather than three days."