Review this: the payments team's primary dashboard has 62 panels on one page — CPU, memory, disk and network for each of 9 services, JVM heap and GC for each, connection-pool gauges, Kafka consumer lag, and three panels at the bottom showing authorisation rate, settlement latency and decline reasons. On-call says it is "comprehensive". The last two incidents were diagnosed from logs. What do you remove, what do you change, and what do you leave alone?
Show the full answer Hide the answer
What is actually required
A dashboard has exactly one job, and it is not completeness. It must let someone woken at 03:00 answer, in under a minute: is the service broken, for whom, and is it getting worse? Everything that does not serve that question is competing with it for attention.
The evidence that this dashboard fails its job is in the stem and nobody has read it: the last two incidents were diagnosed from logs. Sixty-two panels were displayed and none of them was the one that mattered. That is not a gap to be filled with a sixty-third panel.
What I would remove, and why it is safe to
- The per-service resource grid — 36 of the 62 panels. Utilisation is the wrong layer for the first question: a service can be at 20% CPU and completely broken, or at 90% and perfectly healthy. Utilisation belongs on a capacity view somebody reads on a Tuesday afternoon, not on an incident view. Safe to remove because it is not deleted — it moves one click away, where it is more readable for having room.
- The JVM heap and GC panels. These are diagnosis, not detection. They matter once you already suspect a specific service, which is step three, not step one. Move them to a per-service drill-down page.
- Connection-pool gauges as instantaneous values. A gauge sampled every 15 seconds misses the four-second exhaustion that caused every timeout in the incident. If the signal matters it must be a maximum over the interval or a counter of acquisition waits, not a point sample. As drawn it is misleading, which is worse than absent.
The one change that matters
Put the three business panels at the top, make them the largest thing on the page, and express each as a ratio against its normal.
Authorisation rate, settlement latency and decline reasons are the only panels that answer "is the service broken, and for whom". On a payments system they are the service. They are currently at the bottom of a page of sixty-two panels, which means that in practice they do not exist.
Express them comparatively: authorisation rate against the same hour last week, not as an absolute percentage. 94% is bad only if you know today should be 97%, and that depends on the hour, the day and the issuer mix. Last week's line on the same axis makes a 3-point drop visible at a glance.
Then add the one thing missing from all 62: a breakdown of the decline and error rate by dimension — by issuer, by card scheme, by merchant tier, by region. Almost every payments incident is "one of these got worse", and a dashboard that shows only aggregates can never show which. This is also the panel that would have replaced both of those log investigations.
What I would leave, even though it looks odd
Kafka consumer lag. It looks like infrastructure plumbing and belongs on the removal list by the logic above. It is not, on this system. In a payments flow, settlement and reconciliation typically run off the log, so consumer lag is the leading indicator of a business outcome: money that has been authorised and not yet settled. It is the one resource-shaped metric here that converts directly into a customer-visible and finance-visible problem, and it leads the business metric rather than trailing it. Keep it, and relabel it from "consumer lag" to what it means — "unsettled transaction backlog" — so that the person at 03:00 does not have to make the translation.
When not to touch it
If on-call diagnosed the last two incidents from the first panel they looked at, leave the dashboard alone whatever it looks like. Sixty-two panels on a page used by three people who built it and know where everything is costs nothing, and the rebuild costs a week plus the period where nobody can find anything. The trigger for this work is evidence of failure — an incident where the signal was present and not found, a new joiner who could not orient, an alert whose dashboard link led nowhere useful. Reorganising a dashboard because it offends a principle is a way to spend a week and lose goodwill, and it generates a second dashboard rather than replacing the first, because the original owners quietly keep theirs.
How I would argue this in the review
Not as an aesthetic preference, because that argument loses. Take the last two incidents and walk the dashboard against the timeline. For each, ask which panel first showed the problem, and at what minute. If the answer is "none of them, we found it in logs", the dashboard has failed a test it was given twice, and the proposal is a response to evidence rather than taste.
Then make the removal reversible and say so: nothing is deleted, and if on-call misses the grid it comes back in a week. Most resistance is loss aversion and it evaporates when the panels are demonstrably one click away. Finally, agree the rule that stops regrowth: a panel earns the incident view only if someone can name an incident where it would have been the first signal. Dashboards accrete because nothing ever has to justify staying.