Review this incident dashboard. One page holds 28 panels, the default range is 6 hours, auto-refresh is 10 seconds, the scrape interval is 30 seconds, and every panel plots a 5-minute rate over per-pod series for a 900-pod fleet. During the last incident the page took 40 seconds to load and two panels timed out. What would you remove, what would you change, and what would you leave alone?
Show the full answer Hide the answer
What the page is actually for
A triage dashboard answers three questions in the first ninety seconds: is the user affected, which dependency moved, and did we cause it. Everything that does not serve one of those belongs on a linked page. This one is a per-pod exploration surface that has been pressed into service as a triage page, and it fails at both jobs.
The two defects that matter
It is a load generator aimed at the system you need most. 28 panels refreshing every 10 seconds is 2.8 range queries per second from one browser tab. Thirty engineers open it during an incident, so the metrics backend takes roughly 84 range queries per second, each selecting 900 series and materialising 720 points per series at 30-second steps — on the order of 600000 samples per panel per refresh. The query tier saturates, panels time out, and the outage now includes the tool.
Refreshing faster than the scrape interval is free of information and not free of cost. A 10-second refresh on a 30-second scrape re-renders identical data twice out of three times. Setting refresh to the scrape interval cuts query load by two thirds and changes nothing a human sees.
Resolution hides the events worth seeing. A 6-hour range rendered into about 1000 pixels buckets at roughly 20 seconds, and the panel's reducer is avg. A 45-second saturation spike is averaged into one or two buckets and flattened. The rule: choose the range so the bucket width is shorter than the shortest event you must detect, and use max rather than avg on any panel meant to catch spikes.
What I would remove
- Per-pod series from the triage page. Plot fleet quantiles and a count of pods breaching a threshold; per-pod breakdown moves to a drill-down page opened deliberately. This is the change that fixes the query cost, because matched series per query is what the backend pays for.
- The panels nobody reads. Check the dashboard's own view counts if the tool exposes them, then delete. Twenty-eight panels on one page means no one has a mental model of it, so during an incident people scroll instead of reading.
- The 6-hour default. 30 minutes for triage, with the range as a one-click change.
What I would leave alone, though it looks odd
The ingestion-lag panel. It looks like backend trivia next to service metrics, and it is the panel that tells you whether to trust the other 27. When the pipeline falls behind, every other line on the page is describing a few minutes ago, and a decision made on it is a decision made about the past. Keep the deploy annotation rail too: correlating a change against a symptom is the single most-used act of triage.
When this critique is wrong
At four services and fifty pods, per-pod series on one page is the right dashboard and query cost is irrelevant. Density is only a defect once the operator cannot hold the page in their head or the backend cannot serve it. The flip condition is measurable: matched series per refresh and page load time under incident-level concurrency. If the page loads in a second with thirty viewers, leave it alone and spend the effort elsewhere.