concept

Metric Cardinality

The number of unique label combinations on a metric, which determines the number of time series stored and is the main driver of monitoring cost and failure.

observabilitycostmetrics

Every distinct combination of label values creates a separate time series. A metric with a status code label (5 values), a region label (4) and an endpoint label (50) creates 1,000 series, which is fine. Add a user identifier and it creates one series per user, which is not.

The failure is abrupt rather than gradual: monitoring systems degrade badly under cardinality explosion, queries slow, ingestion falls behind, and the observability platform becomes unavailable — frequently during an incident, because incidents are when unusual label values proliferate.

The labels that cause it are consistent and worth prohibiting by convention: user identifiers, request identifiers, session identifiers, email addresses, full URL paths with parameters embedded, timestamps, and error messages used as label values.

The rule that resolves it cleanly is a division of labour between signal types. Metrics are for aggregate, bounded-cardinality questions — how many requests, how slow, what error rate — and their labels must have small, closed value sets. Traces and logs are for high-cardinality investigation: which user, which request, what exact parameters. Trying to answer a per-user question with metrics is what produces the explosion.

The control worth putting in the platform: cardinality limits per metric with alerting before the limit, since detecting this after ingestion has broken is the expensive path, and the offending deployment is usually easy to identify if you are told within minutes.