Term Kind Topic What it is
Alert Fatigue concept Observability The desensitisation that follows from alerts that are frequent, non-actionable, or not tied to user impact — after which real alerts are missed too.
Alert Fatigue in Practice concept Alert Fatigue The state in which alerts are ignored because most of them do not matter — a reliability failure caused by monitoring rather than prevented by it.
Alert on Symptoms Symptom-Based Alerting concept Alerting Paging on user-visible impact rather than on internal conditions that may or may not cause it.
Cardinality concept Observability The number of distinct time series produced by a metric, which is the product of the distinct values of all its labels — and the main driver of monitoring cost.
Counter Reset Handling Monotonic Counter Semantics, Rate Reset Compensation concept Metrics The rule that a decrease in a cumulative counter is interpreted as a process restart rather than as negative work, which is what allows rates to survive deploys and what makes gauges-as-counters silently wrong.
Counter, Gauge and Histogram concept Metrics The three fundamental metric types, distinguished by what they represent over time and by which aggregations are valid.
Distributed Trace concept Distributed Tracing A causally-linked record of one request's path across services, composed of spans that carry timing, attributes and parent relationships.
Exemplar concept OpenTelemetry A trace identifier attached to a metric data point, linking an aggregate measurement directly to a concrete request that produced it.
Liveness vs Readiness concept Health Checks Two distinct questions - should this process be restarted, and should it receive new work - which must be answered by different checks with opposite sensitivities.
Log Schema Drift Field Type Drift, Telemetry Schema Breakage concept Structured Logging The silent breakage of dashboards and alerts when a service changes the name, type or nesting of a structured log field, producing empty results rather than errors.
Metric Cardinality concept Cardinality The number of unique label combinations on a metric, which determines the number of time series stored and is the main driver of monitoring cost and failure.
Metric Staleness Staleness Marker, Stale Series concept Metrics The rules deciding how long a time series keeps answering queries after it stops being written - which determines whether a dead process's last value is served as current, and whether an alert on a vanished se…
Metric Temporality Delta Temporality, Cumulative Temporality concept OpenTelemetry Whether a reported metric point carries a running total since a start time or only the change during one interval - the choice that decides whether process restarts or duplicate deliveries are what corrupts yo…
Metrics concept Metrics Cheap pre-aggregated numeric time series — excellent for knowing something is wrong, structurally unable to tell you which request.
Observability concept Observability The property of being able to answer new questions about a system's internal state from its external outputs, without shipping new code.
Span Link Trace Link, Span Reference concept Distributed Tracing A pointer from one span to another span in a different trace, recording a causal relationship where no single enclosing parent exists - the mechanism that keeps batch jobs, fan-in consumers and delayed asynchr…
Symptom-Based Alerting Alert on Effects, User-Visible Alerting concept Alert Fatigue Paging on what the user experiences and investigating causes with dashboards - because cause-based alerts fire in clusters during a single incident and bury the signal they exist to provide.
Wide Event Canonical Log Line, Arbitrarily-Wide Structured Event, Unit-of-Work Event concept Observability A single structured record per unit of work carrying every field that might matter, aggregated only at query time - so questions nobody anticipated remain answerable, which pre-aggregated metrics make permanen…