Term Kind Topic What it is
Alert Actionability metric Alert Fatigue The proportion of alerts that result in a human taking action, used as the primary quality measure of an alerting system.
Alert Fatigue concept Observability The desensitisation that follows from alerts that are frequent, non-actionable, or not tied to user impact — after which real alerts are missed too.
Alert Fatigue in Practice concept Alert Fatigue The state in which alerts are ignored because most of them do not matter — a reliability failure caused by monitoring rather than prevented by it.
Alert on Symptoms Symptom-Based Alerting concept Alerting Paging on user-visible impact rather than on internal conditions that may or may not cause it.
APM Transaction Tracing tool Application Performance Monitoring Instrumentation that attributes application latency and errors to specific code paths, database queries and external calls, usually with automatic tracing.
Application Performance Monitoring APM tool Application Performance Monitoring Instrumentation that attributes latency and errors to code paths and dependencies inside a service — the layer between metrics and a profiler.
Application Performance Monitoring APM tool Observability Instrumentation inside the application that attributes latency and errors to specific code paths, queries and dependencies.
Burn Rate Alerting Multi-Window Multi-Burn-Rate, SLO Burn Alert, Budget Consumption Alert pattern Alerting Alerting on how fast an error budget is being consumed, across multiple time windows simultaneously, so that a sudden outage pages immediately while a slow erosion opens a ticket.
Business Correlation Identifier Domain Trace Key, Entity Correlation practice Correlation IDs A domain identifier - shipment, order, trip, claim - carried on every log line, event, job and external call, answering the question distributed tracing cannot: what happened to this thing, over days, across s…
Business Metric Alerting Outcome Monitoring practice Business Metrics Alerting on product outcomes rather than on technical signals, because it is the only way to catch failures where the system works correctly and produces the wrong result.
Business Metric Instrumentation practice Business Metrics Emitting metrics for business outcomes — orders, payments, signups — alongside technical telemetry, so incidents can be detected and prioritised by impact.
Business Metrics metric Business Metrics Measuring what the system exists to do — frequently the fastest and most reliable outage detector available.
Cardinality concept Observability The number of distinct time series produced by a metric, which is the product of the distinct values of all its labels — and the main driver of monitoring cost.
Cardinality Budget Series Budget, Label Quota practice Cardinality A measured, owned limit on the number of distinct metric series a team may create - the control that converts an invisible cost externality into a visible constraint.
Context Propagation pattern Correlation IDs Carrying request-scoped identifiers and metadata across every service, thread and asynchronous boundary so a single flow remains traceable end to end.
Continuous Profiling practice Profiling Sampling CPU, memory and lock profiles from production continuously at low overhead, so resource usage can be attributed to specific code paths.
Correlation ID practice Observability A single identifier attached to one logical operation and included in every log line it produces, anywhere in the system.
Correlation IDs Request ID, Trace ID pattern Correlation IDs A single identifier attached at the edge and carried through every hop, synchronous and asynchronous, that stitches an investigation together.
Counter Reset Handling Monotonic Counter Semantics, Rate Reset Compensation concept Metrics The rule that a decrease in a cumulative counter is interpreted as a process restart rather than as negative work, which is what allows rates to survive deploys and what makes gauges-as-counters silently wrong.
Counter, Gauge and Histogram concept Metrics The three fundamental metric types, distinguished by what they represent over time and by which aggregations are valid.
Critical Path Analysis Span Critical Path, Blocking Time Analysis practice Distributed Tracing Identifying which spans in a trace actually block the response, as opposed to running in parallel, so optimisation effort targets the work that determines latency rather than the work that merely appears slow.
Cross-Service Debugging practice Debugging Distributed Systems Investigating a failure that spans multiple services by moving between traces, logs, metrics and profiles along a single correlated request.
Dashboard Hierarchy practice Dashboards Organising dashboards into service-level, diagnostic and deep-dive layers so each answers a specific question at a specific moment.
Dashboards practice Dashboards Curated views built for a specific question — with the failure mode of showing everything and answering nothing.
Debugging Distributed Systems practice Debugging Distributed Systems A method for diagnosing failures whose cause is in a different service from the symptom.
Differential Profiling Profile Diffing, Comparative Profiling practice Profiling Comparing two normalised profiles of one workload to attribute a performance change to a specific code path - the step that turns "release 2.4 is slower" into a named function, and the normalisation errors tha…
Distributed Trace concept Distributed Tracing A causally-linked record of one request's path across services, composed of spans that carry timing, attributes and parent relationships.
Distributed Tracing tool Observability Following one logical request across every service it touches by propagating a shared trace identifier and recording timed spans.
Distributed Tracing in Practice tool Distributed Tracing Following one request across every service it touches — the only tool that answers "where did the time go" in a distributed system.
Error Fingerprinting Issue Grouping, Event Deduplication pattern Observability Deriving a stable identifier from an error's invariant attributes so that many occurrences collapse into one actionable issue - and the two directions in which the heuristic fails damagingly.
Exemplar concept OpenTelemetry A trace identifier attached to a metric data point, linking an aggregate measurement directly to a concrete request that produced it.
Golden Signals metric Observability The four measurements that cover most of what matters for a request-driven service: latency, traffic, errors and saturation.
Health Check practice Observability An endpoint the platform polls to decide whether an instance should be restarted or should receive traffic — two different questions needing two different checks.
Health Check Semantics practice Health Checks The distinction between liveness, readiness and startup checks, and the failure each is intended to address.
Health Checks pattern Health Checks Endpoints that tell the platform whether to route traffic to an instance or restart it — and a well-documented way to amplify an outage.
Ingestion-Query Isolation Write-Read Workload Separation, Analytics Isolation pattern Observability Separating a continuous high-volume write path from a spiky, arbitrary query path so that an expensive query cannot stall ingestion - because a query delay is recoverable and an ingestion gap is not.
Liveness vs Readiness concept Health Checks Two distinct questions - should this process be restarted, and should it receive new work - which must be answered by different checks with opposite sensitivities.
Log Level Discipline practice Logging Consistent semantics for log severity so that levels can be used for routing, alerting and cost control.
Log Management practice Log Management The pipeline, storage, retention and access model for logs — where cost and compliance meet operational need.
Log Retention Tiering practice Log Management Storing log data at different resolutions, costs and access latencies according to how old it is and how likely it is to be queried.
Log Sampling Budget Log Volume Budget, Log Ingest Quota practice Logging An owned per-service allowance of log volume, spent by deciding which record classes are kept whole and which are sampled by trace - the control that turns an invisible shared cost into a decision a team makes…
Log Schema Consistency practice Structured Logging Enforcing the same field names, types and semantics for structured log records across every service, so cross-service queries are possible.
Log Schema Drift Field Type Drift, Telemetry Schema Breakage concept Structured Logging The silent breakage of dashboards and alerts when a service changes the name, type or nesting of a structured log field, producing empty results rather than errors.
Logging practice Logging Emitting events a human or a query can reason about later — where the discipline is structure, sampling and what you refuse to log.
Metric Cardinality concept Cardinality The number of unique label combinations on a metric, which determines the number of time series stored and is the main driver of monitoring cost and failure.
Metric Staleness Staleness Marker, Stale Series concept Metrics The rules deciding how long a time series keeps answering queries after it stops being written - which determines whether a dead process's last value is served as current, and whether an alert on a vanished se…
Metric Temporality Delta Temporality, Cumulative Temporality concept OpenTelemetry Whether a reported metric point carries a running total since a start time or only the change during one interval - the choice that decides whether process restarts or duplicate deliveries are what corrupts yo…
Metrics concept Metrics Cheap pre-aggregated numeric time series — excellent for knowing something is wrong, structurally unable to tell you which request.
Multi-Window Multi-Burn-Rate Alerting practice SLO Monitoring Alerting on how fast an error budget is being consumed, over two time windows simultaneously, to get both fast detection and few false alarms.
Observability concept Observability The property of being able to answer new questions about a system's internal state from its external outputs, without shipping new code.
Observability in Practice Telemetry, Instrumentation pattern Observability The three signals, what each is actually for, and why the links between them matter more than any of them individually.
OpenTelemetry OTel tool OpenTelemetry A vendor-neutral standard and toolset for generating, collecting and exporting traces, metrics and logs.
Outlier Alerting Per-Tenant Alerting, Worst-Case Alerting practice Application Performance Monitoring Alerting on the worst-affected segment rather than on an aggregate, because in skewed populations the aggregate is dominated by the segment that is fine.
Page and Ticket Routing practice Alerting Deciding for each monitored condition whether it warrants immediate human interruption or asynchronous handling, and enforcing the distinction.
Profiling practice Profiling Attributing resource consumption to specific code paths — the tool for "why is this slow" once you know where.
RED Method practice Observability A minimal per-service dashboard: Rate, Errors, Duration — the request-centric view of whether users are being served.
Runbook Playbook practice Observability A short, actionable document telling an on-call engineer what an alert means, what to check, and what the safe mitigations are.
SLI Measurement Point SLI Vantage Point, Indicator Measurement Location practice SLO Monitoring The deliberate choice of where in the request path an indicator is computed, which determines which failures the objective can see at all.
Span Attribute Budget Trace Attribute Discipline, Span Cardinality Budget practice Distributed Tracing A deliberate limit on how many spans a request emits and how many attributes each carries, set so that traces stay readable and affordable rather than complete.
Span Link Trace Link, Span Reference concept Distributed Tracing A pointer from one span to another span in a different trace, recording a causal relationship where no single enclosing parent exists - the mechanism that keeps batch jobs, fan-in consumers and delayed asynchr…