Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
4 to work through
-
beginner
Your team's on-call phone has not been paged in nine days and every dashboard is green. Why is that not yet evidence that anything is healthy, and what single mechanism turns silence into a signal?
2 min answer -
intermediate
A platform's alerts are defined per component - CPU, memory, queue depth, replica lag. During an incident, forty alerts fire simultaneously. What should the alerting strategy be instead?
2 min answer -
intermediate
What is the test for whether an alert should exist, and what does applying it honestly do?
2 min answer -
advanced
Error rate has been 1.5% for three days. No alert fired. Customers are complaining. What is happening and what failed?
2 min answer
3 terms in this topic
Alert on Symptoms
Paging on user-visible impact rather than on internal conditions that may or may not cause it.
patternBurn Rate Alerting
Alerting on how fast an error budget is being consumed, across multiple time windows simultaneously, so that a sudden outage pages immediately while …
practicePage and Ticket Routing
Deciding for each monitored condition whether it warrants immediate human interruption or asynchronous handling, and enforcing the distinction.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.