Your metrics platform ingests roughly 40 million samples per second and holds about 3 billion active series, the scale eBay published for Sherlock.io in 2022. A team asks to add a `customer_id` label to the request-latency histogram so they can answer per-customer questions. Estimate what that costs, and say what you would offer them instead.
Show the full answer Hide the answer
The assumptions, stated
Four numbers decide this, and only one of them is in the request.
- The histogram is not one series. A Prometheus-style histogram with 12 buckets emits 12
_bucketseries plus_sumand_count— 14 series per label combination, not one. - The existing label set already multiplies. Say the metric carries
service,endpoint,method,statusandpod. If the team's service has 40 endpoints, 4 methods, 6 status classes and 300 pods, that is 40 × 4 × 6 × 300 ≈ 288,000 combinations beforecustomer_id. customer_idis a tenant identifier. On a marketplace that is not 50 values. Take 20,000 active business customers in any scrape window as a deliberately conservative figure.- Series cost is roughly 1–3 KB of resident memory each in the Prometheus head block, plus index. Take 2 KB as the planning number.
The arithmetic, shown
The multiplication is the whole answer:
288,000 combinations × 14 histogram series = 4,032,000 series today
× 20,000 customer_id values ≈ 80,640,000,000 series
Eighty billion series against a platform that holds three billion in total. At 2 KB each that is on the order of 160 TB of head memory for one metric on one service. The request is not expensive; it is arithmetically impossible, and it is impossible by a factor of roughly 27,000.
Clamp every assumption to the friendliest end — 6 buckets, 2,000 customers, one pod — and it is still 40 × 4 × 6 × 2,000 × 8 ≈ 15 million series for one service's latency metric, half a percent of the entire platform spent on one team's drill-down.
Which assumption dominates the error
Not customer_id. It is pod. Pod identity is already in the label set on most Kubernetes deployments, it changes on every deploy, and it multiplies everything to its right. Dropping pod from this metric removes a 300x factor and makes the churn problem — series that exist for twenty minutes and are retained for the full block — go away with it. Before refusing the new label, audit the one that is already there and earning nothing, because nobody dashboards per-pod latency; they dashboard per-service latency and go to logs for the pod.
Active series are also index entries every query must walk, so one team's cardinality degrades the p99 of the query the on-call engineer runs at 03:00.
What the number rules in, and what to offer instead
The estimate rules out the label, and it rules in three cheaper answers:
- Exemplars. Attach a sampled trace ID to the histogram bucket rather than a label to the series. The team clicks from the slow bucket into an actual trace for an actual customer. Cardinality cost: zero new series. This answers "which customer is slow" without answering it for all customers at once, which is what they actually need at 03:00.
- Wide events for the drill-down. One record per request, with
customer_idas a column rather than a series, queried at read time. A column with 20,000 distinct values costs the same as a column with three. This is the structurally correct home for per-customer questions. - A tiered metric. Keep the histogram unlabelled by customer and emit a second, narrow metric — a counter and a
_sum, no buckets, two series each — for the 50 contract customers with a latency SLA. 100 series buys the per-customer answer for the customers it is contractually owed to. This is almost always the real requirement, stated badly.
When the refusal is wrong
If the service genuinely has a handful of tenants — an internal platform with 30 consumers, a B2B product with 80 accounts — then customer_id is a small label and refusing it is cargo-culted caution. 80 × 14 × a modest base set is tens of thousands of series, which is noise at this scale. The rule is not "never add a tenant label"; it is "multiply it out against the existing label set before you answer." A label is cheap or ruinous depending entirely on what it multiplies.
Common weak answers
- "Add it and set a retention policy." Retention governs how long blocks are kept on disk. Active series are a live memory and index cost in the current head block. Shorter retention does not reduce the peak by a single byte.
- "Use a recording rule." Recording rules aggregate an existing series into a new one. The expensive series must exist first for the rule to read it, so the rule adds cost rather than removing it.
- Quoting a storage-cost-per-GB figure. The binding constraint is head memory and index width, not object storage, which is cheap. Answering in dollars per terabyte misses the failure mode, which is the metrics tier going OOM.