Dashboards
Answering 'is it us' in under a minute, for someone who was asleep.
4 to work through
-
intermediate
A platform has hundreds of dashboards and engineers cannot find the right one during an incident. How should dashboards be structured?
2 min answer -
intermediate
Design the one dashboard your team opens first during an incident. What is on it and what is deliberately not?
2 min answer -
intermediate
Review this incident dashboard. One page holds 28 panels, the default range is 6 hours, auto-refresh is 10 seconds, the scrape interval is 30 seconds, and every panel plots a 5-minute rate over per-pod series for a 900-pod fleet. During the last incident the page took 40 seconds to load and two panels timed out. What would you remove, what would you change, and what would you leave alone?
3 min answer -
intermediate
Your metrics pipeline normally makes data queryable about 20 seconds after it is emitted. During a large incident the ingestion tier falls behind and the lag grows to four minutes. Nothing is lost. Walk through what happens to the on-call engineer's decisions.
3 min answer
2 terms in this topic
Neighbouring topics
Observability
General material on understanding a system from its outputs.
Logging
What to log, at what level, and what must never appear in a log.
Structured Logging
Machine-parseable events with stable names and consistent fields.
Metrics
Counters, gauges and histograms, and percentiles rather than means.
Cardinality
The label that multiplies series count and the bill with it.
Distributed Tracing
Reconstructing one request's path across every service it touched.
Correlation IDs
One identifier propagated through every hop and every log line.
Sampling
Head-based versus tail-based, and keeping the traces that matter.
Health Checks
Liveness versus readiness, and the check that causes the outage.
Alerting
Symptom-based, actionable, user-impacting — and linked to a runbook.
Alert Fatigue
How noise makes the real page invisible, and the structural fix.
Application Performance Monitoring
Attributing latency to code paths, queries and dependencies.
Profiling
Continuous CPU and memory attribution in production.
Business Metrics
Orders per minute alongside error rate, because healthy is not enough.
SLO Monitoring
Burn-rate alerting that fires on user impact rather than on thresholds.
Log Management
Aggregation, retention tiering, search and the cost of keeping everything.
Telemetry Cost
Observability bills that rival compute, and where to cut without going blind.
Debugging Distributed Systems
Localising a regression when every service reports healthy.
OpenTelemetry
Instrumenting once against an open standard rather than a vendor agent.