metric

Active Series

also called Live Time Series, Head Series

The count of distinct label combinations currently receiving samples, which is the quantity a metrics tier's memory and query latency actually scale with - not the volume of samples ingested.

cardinalityprometheustelemetry costcapacityebay

A team asks to add customer_id to a latency histogram. The request sounds like one field. On a platform at eBay's published Sherlock.io scale — roughly 1.5 million Prometheus endpoints scraped every minute, about 40 million samples per second, 3 billion active series (eBay engineers, 2022) — it is arithmetically impossible, and nobody in the conversation can see why without doing the multiplication.

An active series is one unique combination of metric name and label values that is currently receiving data. http_latency{service="checkout", endpoint="/pay", status="200", pod="abc"} is one series. Change any label value and it is a different series, with its own memory, its own index entry and its own chunk on disk.

The number that governs a metrics platform is active series, not samples per second. Sample rate scales the write path, which is cheap and linear. Series count scales resident memory, index width and query fan-out, which is expensive and is where the platform falls over.

Why it matters

Series count is multiplicative across labels. Forty endpoints, four methods, six status classes and 300 pods is 288,000 combinations before anything interesting is added. Multiply by 20,000 customers and the result is in the tens of billions — not a larger bill, a different order of reality.

Two costs follow, and the second is the one that surprises people:

  • Memory. Roughly 1–3 KB of resident memory per active series in the head block, plus index overhead. Take 2 KB as a planning figure. A million series is a couple of gigabytes; a billion is not survivable on one machine.
  • Query latency for everyone. Series are index entries, and every query walks the index for matching series before it reads a single sample. One team's cardinality raises the p99 of the query the on-call engineer runs at 03:00, on an unrelated service. The cost is externalised, which is exactly why it needs a budget rather than good intentions.

Implementation patterns

  • Multiply before approving. The review question is never "is this label big?" but "what does it multiply?" A label with 50 values on a metric with 200,000 existing combinations is 10 million new series.
  • Count the histogram. A histogram is one series per bucket plus _sum and _count. Twelve buckets is 14 series per label combination. This factor is omitted from almost every informal estimate.
  • Audit the labels already there. pod is usually the worst offender and nobody dashboards it: it multiplies everything to its right, and it churns on every deploy, leaving dead series retained for the rest of the block.
  • Per-team series budgets, enforced at the collector. A relabelling rule that drops unapproved high-cardinality labels at ingest, with a dashboard of each team's consumption. Budgets without enforcement are advisory and are ignored.
  • Push identity out of labels. Exemplars attach a sampled trace ID to a bucket at zero series cost. Wide events put customer_id in a column, where 20,000 distinct values cost the same as three.
  • Tier deliberately. Aggregate metrics for everyone, plus a narrow per-tenant counter for the 50 customers with a contractual SLA. 100 series buys the answer that actually had to be bought.

Industry example

eBay's move to OpenTelemetry, published in 2022, describes a metrics tier holding about 3 billion active series fed by roughly 40 million samples per second from 1.5 million scraped endpoints. The useful thing in those numbers is their ratio: there are far more active series than samples per second, which means the typical series is scraped infrequently and still costs memory every moment it exists. A platform at that scale is not bounded by how fast it can ingest. It is bounded by how many distinct things it is asked to remember.

Failure scenarios

  • Cardinality explosion from an unbounded label. Someone adds a user ID, a URL path with IDs in it, or an error message as a label. Series count rises by orders of magnitude within one deploy and the metrics tier OOMs.
  • Churn without growth. Pod names change on every deploy, so a steady-state fleet generates new series continuously. The live count looks stable while the block's index grows all day, and the tier dies at retention boundaries rather than at deploys, which makes the cause hard to see.
  • Silent query failure. Many systems cap series per query. Past the cap, the dashboard silently shows a subset, and the graph looks plausible while being wrong — the worst outcome, since nobody is alerted to it.
  • Cross-team blast radius. One team's explosion degrades every other team's queries and alert evaluations, so alerts for unrelated services start evaluating late or timing out.

Trade-offs

Choose Gains Pays
Rich labels on metrics Drill-down without leaving the metrics tool Multiplicative memory and index cost, borne by every tenant
Aggregate metrics plus exemplars Near-zero series cost, click through to a real trace Only sampled requests are reachable
Wide events for drill-down Cardinality is free; any field is queryable at read time A second system, shorter retention, higher per-byte cost

When not to use it

Series counting is not the right lens for a small deployment. A team with 20 services and no multi-tenant dimension has tens of thousands of series, and the machinery above — budgets, relabelling rules, enforcement — is governance overhead against a cost that would fit on one modest instance. Add the label, watch the count monthly, and spend the attention elsewhere. The lens becomes essential at the point where a label's cardinality is set by customer count rather than by engineering decisions, because that is where growth stops being under your control.

Equally, do not use series count as a proxy for telemetry spend if the vendor bills on ingested volume or on hosts: optimise the thing you are charged for. Series count is the right lens for whether the platform will stay up.

Interview question

Q: A team wants per-customer latency percentiles for 20,000 customers. Walk me through how you would answer them.

What a strong answer covers: doing the multiplication aloud, including the histogram's bucket factor · identifying that the existing pod label is probably the larger problem · distinguishing memory cost from query-latency cost borne by other teams · offering exemplars, wide events, or a narrow per-tenant metric for the subset under an SLA · and recognising that "20,000 customers need percentiles" is usually "50 contract customers need percentiles", stated imprecisely.

Quick check

Quiz: A metric has 5,000 label combinations and uses a 10-bucket histogram. How many active series? — 5,000 × 12 = 60,000 (ten buckets plus _sum and _count).

Flashcard: What does a metrics platform's memory actually scale with? — Active series, not samples per second. Roughly 1–3 KB resident each, and series multiply across labels while samples merely add.