SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
4 to work through
-
intermediate Multiple choice
A checkout API has a 99.9% availability SLO and the team must decide where the indicator is computed from. The candidates are load-balancer access logs, in-process server metrics, the mobile client's own reporting, and synthetic probes. Which should be the primary source?
3 min answer -
intermediate Multiple choice
A food-delivery platform in Zomato's mould pushes order-status events to restaurant tablets and to customers through a queue-backed webhook fleet. The complaints are that status arrives late rather than that it never arrives. Which indicator should the SLO be written on?
3 min answer -
advanced
A canary release looks healthy on p50 latency but a small set of enterprise tenants sees timeouts. Which metrics and gates should have caught it?
2 min answer -
advanced
A communication platform sets a 99.9% availability SLO. How should alerting on that SLO be structured so it catches both sudden outages and slow degradation?
2 min answer
2 terms in this topic
Multi-Window Multi-Burn-Rate Alerting
Alerting on how fast an error budget is being consumed, over two time windows simultaneously, to get both fast detection and few false alarms.
practiceSLI Measurement Point
The deliberate choice of where in the request path an indicator is computed, which determines which failures the objective can see at all.
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.