Cardinality
The label that multiplies series count and the bill with it.
5 to work through
-
advanced Multiple choice
A live-streaming platform of Twitch's shape needs usable latency quantiles for 300 API endpoints whose responses span 2 ms to 90 seconds. The current histograms use 12 fixed buckets topping out at 10 seconds and every p99 above that reads as the overflow bucket. Which change fits the problem?
3 min answer -
advanced
A single deploy took down your monitoring platform. What happened, and how do you prevent a recurrence?
2 min answer -
advanced
A team wants to debug failures nobody predicted, using wide high-cardinality events rather than pre-aggregated metrics. How does the storage design differ, and how must sampling preserve rare errors?
3 min answer -
advanced
An observability platform ingests billions of telemetry events daily, and some customers attach dimensions with unbounded distinct values. How should ingestion, storage, indexing and query be designed so cost and latency stay manageable?
2 min answer -
advanced
An observability platform ingests enormous telemetry volume, but a small number of dimensions create extreme cardinality. How should ingestion, aggregation, indexing, sampling, retention and storage tiers be designed so cost and query performance stay predictable?
2 min answer
2 terms in this topic
Cardinality Budget
A measured, owned limit on the number of distinct metric series a team may create - the control that converts an invisible cost externality into a vi…
conceptMetric Cardinality
The number of unique label combinations on a metric, which determines the number of time series stored and is the main driver of monitoring cost and …
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.