Observability
General material on understanding a system from its outputs.
8 to work through
-
intermediate
A new service goes live next week. What observability must exist on day one?
2 min answer -
intermediate
You are designing a new service. What must be in place before it goes live so that whoever is paged at 3 AM can diagnose it without you?
2 min answer -
advanced
A commerce platform can tell that its peak-day checkout success rate has dropped, but cannot tell which of forty services is responsible. What is missing, and in what order should it be added?
2 min answer -
advanced
An analytics platform ingests billions of events while customers run arbitrary segmentation queries. How should the two workloads be isolated, and when should pre-aggregation be introduced?
2 min answer -
advanced
An error-tracking platform receives millions of events, many of them the same underlying problem with different stack details. How should grouping, deduplication, high-cardinality metadata and retention be designed?
2 min answer -
advanced
On 11 December 2024 OpenAI rolled a new telemetry service out across every Kubernetes cluster. Within about half an hour the API servers were saturated and services could no longer resolve one another; full recovery took until the evening. What turned a monitoring change into a total outage, and which of the contributing factors would you fix first?
3 min answer -
advanced
Your observability bill is now 40% of your compute bill. Leadership wants it cut in half without going blind. What do you cut?
2 min answer -
advanced
p99 latency on checkout tripled overnight. Dashboards look normal, no deployment went out, and every service reports healthy. How do you find it?
3 min answer
17 terms in this topic
Alert Fatigue
The desensitisation that follows from alerts that are frequent, non-actionable, or not tied to user impact — after which real alerts are missed too.
toolApplication Performance Monitoring
Instrumentation inside the application that attributes latency and errors to specific code paths, queries and dependencies.
conceptCardinality
The number of distinct time series produced by a metric, which is the product of the distinct values of all its labels — and the main driver of monit…
practiceCorrelation ID
A single identifier attached to one logical operation and included in every log line it produces, anywhere in the system.
toolDistributed Tracing
Following one logical request across every service it touches by propagating a shared trace identifier and recording timed spans.
patternError Fingerprinting
Deriving a stable identifier from an error's invariant attributes so that many occurrences collapse into one actionable issue - and the two direction…
metricGolden Signals
The four measurements that cover most of what matters for a request-driven service: latency, traffic, errors and saturation.
practiceHealth Check
An endpoint the platform polls to decide whether an instance should be restarted or should receive traffic — two different questions needing two diff…
patternIngestion-Query Isolation
Separating a continuous high-volume write path from a spiky, arbitrary query path so that an expensive query cannot stall ingestion - because a query…
conceptObservability
The property of being able to answer new questions about a system's internal state from its external outputs, without shipping new code.
patternObservability in Practice
The three signals, what each is actually for, and why the links between them matter more than any of them individually.
practiceRED Method
A minimal per-service dashboard: Rate, Errors, Duration — the request-centric view of whether users are being served.
practiceRunbook
A short, actionable document telling an on-call engineer what an alert means, what to check, and what the safe mitigations are.
metricTelemetry Pipeline Lag
The delay between an event being emitted and being queryable, which bounds how quickly any decision made from telemetry can respond to reality.
practiceTelemetry Sampling
Keeping a subset of traces or events to bound observability cost, chosen so the ones that matter survive.
practiceUSE Method
For every resource, track Utilisation, Saturation and Errors — the resource-centric complement to request-centric monitoring.
conceptWide Event
A single structured record per unit of work carrying every field that might matter, aggregated only at query time - so questions nobody anticipated r…
1 artifact you would hand over
Neighbouring topics
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.