Term Kind Topic What it is
Business Correlation Identifier Domain Trace Key, Entity Correlation practice Correlation IDs A domain identifier - shipment, order, trip, claim - carried on every log line, event, job and external call, answering the question distributed tracing cannot: what happened to this thing, over days, across s…
Business Metric Alerting Outcome Monitoring practice Business Metrics Alerting on product outcomes rather than on technical signals, because it is the only way to catch failures where the system works correctly and produces the wrong result.
Business Metric Instrumentation practice Business Metrics Emitting metrics for business outcomes — orders, payments, signups — alongside technical telemetry, so incidents can be detected and prioritised by impact.
Cardinality Budget Series Budget, Label Quota practice Cardinality A measured, owned limit on the number of distinct metric series a team may create - the control that converts an invisible cost externality into a visible constraint.
Continuous Profiling practice Profiling Sampling CPU, memory and lock profiles from production continuously at low overhead, so resource usage can be attributed to specific code paths.
Correlation ID practice Observability A single identifier attached to one logical operation and included in every log line it produces, anywhere in the system.
Critical Path Analysis Span Critical Path, Blocking Time Analysis practice Distributed Tracing Identifying which spans in a trace actually block the response, as opposed to running in parallel, so optimisation effort targets the work that determines latency rather than the work that merely appears slow.
Cross-Service Debugging practice Debugging Distributed Systems Investigating a failure that spans multiple services by moving between traces, logs, metrics and profiles along a single correlated request.
Dashboard Hierarchy practice Dashboards Organising dashboards into service-level, diagnostic and deep-dive layers so each answers a specific question at a specific moment.
Dashboards practice Dashboards Curated views built for a specific question — with the failure mode of showing everything and answering nothing.
Debugging Distributed Systems practice Debugging Distributed Systems A method for diagnosing failures whose cause is in a different service from the symptom.
Differential Profiling Profile Diffing, Comparative Profiling practice Profiling Comparing two normalised profiles of one workload to attribute a performance change to a specific code path - the step that turns "release 2.4 is slower" into a named function, and the normalisation errors tha…
Health Check practice Observability An endpoint the platform polls to decide whether an instance should be restarted or should receive traffic — two different questions needing two different checks.
Health Check Semantics practice Health Checks The distinction between liveness, readiness and startup checks, and the failure each is intended to address.
Log Level Discipline practice Logging Consistent semantics for log severity so that levels can be used for routing, alerting and cost control.
Log Management practice Log Management The pipeline, storage, retention and access model for logs — where cost and compliance meet operational need.
Log Retention Tiering practice Log Management Storing log data at different resolutions, costs and access latencies according to how old it is and how likely it is to be queried.
Log Sampling Budget Log Volume Budget, Log Ingest Quota practice Logging An owned per-service allowance of log volume, spent by deciding which record classes are kept whole and which are sampled by trace - the control that turns an invisible shared cost into a decision a team makes…
Log Schema Consistency practice Structured Logging Enforcing the same field names, types and semantics for structured log records across every service, so cross-service queries are possible.
Logging practice Logging Emitting events a human or a query can reason about later — where the discipline is structure, sampling and what you refuse to log.
Multi-Window Multi-Burn-Rate Alerting practice SLO Monitoring Alerting on how fast an error budget is being consumed, over two time windows simultaneously, to get both fast detection and few false alarms.
Outlier Alerting Per-Tenant Alerting, Worst-Case Alerting practice Application Performance Monitoring Alerting on the worst-affected segment rather than on an aggregate, because in skewed populations the aggregate is dominated by the segment that is fine.
Page and Ticket Routing practice Alerting Deciding for each monitored condition whether it warrants immediate human interruption or asynchronous handling, and enforcing the distinction.
Profiling practice Profiling Attributing resource consumption to specific code paths — the tool for "why is this slow" once you know where.
RED Method practice Observability A minimal per-service dashboard: Rate, Errors, Duration — the request-centric view of whether users are being served.
Runbook Playbook practice Observability A short, actionable document telling an on-call engineer what an alert means, what to check, and what the safe mitigations are.
SLI Measurement Point SLI Vantage Point, Indicator Measurement Location practice SLO Monitoring The deliberate choice of where in the request path an indicator is computed, which determines which failures the objective can see at all.
Span Attribute Budget Trace Attribute Discipline, Span Cardinality Budget practice Distributed Tracing A deliberate limit on how many spans a request emits and how many attributes each carries, set so that traces stay readable and affordable rather than complete.
Structured Logging practice Structured Logging Emitting logs as machine-parseable key-value records rather than formatted prose, so they can be queried, aggregated and correlated.
Tail-Based Sampling Post-Trace Sampling, Intelligent Sampling practice Sampling Deciding whether to keep a trace after it completes, so that errors and slow requests are always retained while ordinary traffic is sampled cheaply.
Telemetry Cost Control practice Telemetry Cost Managing the spend on logs, metrics and traces, which in mature estates can approach or exceed the cost of the infrastructure being observed.
Telemetry Cost Management practice Telemetry Cost Controlling observability spend through sampling, retention tiering and cardinality limits without losing diagnostic capability.
Telemetry Sampling practice Observability Keeping a subset of traces or events to bound observability cost, chosen so the ones that matter survive.
Trace Sampling Strategy Head Sampling, Tail Sampling practice Sampling Deciding which traces to keep, since retaining all of them at scale is unaffordable and retaining a random few loses exactly the interesting ones.
USE Method practice Observability For every resource, track Utilisation, Saturation and Errors — the resource-centric complement to request-centric monitoring.