advanced 3 min answer

A platform injects standard labels into every metric it scrapes: team, service, environment, pod and commit SHA. Between deploys the metrics backend is healthy. Within minutes of each fleet deploy, query latency triples, ingester memory climbs and dashboards covering the last six hours time out. Samples per second have not changed. Where is the time going?

prometheuscardinalityseries churnlabelsplatform telemetry
Show the full answer Hide the answer

The first three things I would look at

  1. Active series and new series per minute, not samples per second. The stem tells you ingestion volume is flat, which eliminates the obvious explanation and points at series identity.
  2. Series matched per query for one slow dashboard panel, over a window that spans a deploy versus one that does not.
  3. The label sets of two series for the same metric and pod before and after the deploy, to confirm which labels changed value.

The diagnosis

A series is identified by its full label set, so a label that changes on every deploy does not update a series — it retires one and creates another. With commit SHA and pod name in every series, a fleet deploy replaces the entire active set. Two thousand pods at 150 series each is about 300000 series; three deploys a day create and abandon that many again each time.

Two costs follow, and neither is sample volume:

  • Index and per-series overhead. A query selecting by service over six hours must resolve every series that existed in that window. Spanning three deploys means matching roughly three times the series for the same amount of data, and the index work is per series, not per sample. That is the tripled query latency.
  • Head-block memory. The ingester holds recent series in memory; churn inflates that set well beyond the number of series actually being written now, which is the memory climb that recovers slowly after each deploy.

The arithmetic that shows sample volume is a red herring: Prometheus documents storage as retention times ingest rate times bytes per sample, and states it averages only 1 to 2 bytes per sample. At a 15-second scrape, 300000 series is about 20000 samples per second, or roughly 3.5 GB a day of samples. Cheap. The cost is in identity, not in data.

The misleading signal and why it misleads

Samples per second being flat is exactly the signal that sends teams to the wrong place — usually to query optimisation or a bigger ingester, both of which buy a little headroom and leave the mechanism untouched. Scaling storage for a cardinality problem works until the next fleet deploy.

The fix

  • Take the SHA out of every series. Publish it once per pod as an info metric — one series carrying build metadata — and join it at query time when you need to compare releases. One series per pod rather than one per metric per pod.
  • Drop pod from aggregated series and keep it only where per-instance detail is actually used. Pod names change on every restart, so they are the second churn source.
  • Enforce the label set at the platform boundary. Since the platform injects labels for everyone, it is also the only place a limit can be applied: an allow-list in the scrape configuration, plus a per-target series cap that fails loudly for one team rather than degrading the backend for all of them.
  • Make churn a platform SLO. New series per minute, per team, with a published budget. It is the number that predicts the next outage, and it is not on any default dashboard.

Common weak answers

"Too much cardinality" without separating count from churn leads teams to delete useful dimensions such as service or endpoint while leaving the label that actually churns. "Sample more slowly" reduces samples and not series. "Add ingesters" is the answer that works for three months. The rule: a label whose value changes on deploy or restart belongs in an info metric or an exemplar, never in the identity of every series.