Review this platform telemetry programme. 14 dashboards, a weekly adoption report by capability, per-capability request counters with caller identity, a quarterly developer survey of 38 questions, and DORA metrics computed per product team and published weekly to leadership. Two of the seven platform engineers spend about a day a week each maintaining it. What would you remove, what would you change, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
A platform's telemetry exists to answer three questions: which capabilities are used and by whom, where teams get stuck, and whether the platform's own promises are being met. Everything in this list should be justified against one of those, and the programme costs about 2 engineer-days a week out of 35, roughly 6% of the team, which is a real price for evidence.
What I would remove
Per-team DORA metrics published to leadership. Deploy frequency and lead time are useful to a team reading its own trend. Published upwards per team they become a performance ranking, and the measured behaviour changes: changes are split to raise deploy counts, work is pushed through as hotfixes to shorten lead time, and within two quarters the number describes reporting behaviour rather than delivery. Keep DORA aggregated across the estate as a trend the platform is trying to move, with no team-level league table.
The second removal is the 38-question quarterly survey. Long surveys at low frequency give a response rate that falls every round and data that is already stale when it lands. Five questions monthly, same wording every time, gives a usable trend. Cut the 14 dashboards to the two somebody actually opens - one per- capability usage view, one platform-SLO view - and delete the rest rather than letting them rot into mistrust.
The one change that matters
Nothing in the list says where a team failed. Add a per-capability first-run funnel: for each capability, per consumer, the step at which an attempt stopped and whether it ever completed. Usage counters show who succeeded; the funnel shows the 40% who started provisioning a database on Tuesday and were never seen again. That is the only signal in the set that generates roadmap items by itself, and it is the signal almost every platform telemetry programme is missing.
What I would leave alone
The per-capability counters with caller identity, even though they look like the most boring item. They are the precondition for every deprecation the platform will ever run, for change-impact analysis before a breaking release, and for answering "who would this affect" in minutes rather than weeks. Keep the weekly adoption report too, but report voluntary adoption separately from mandated usage, since only the first tells you anything about quality.
How I would argue this in the review
Ask for the last three roadmap decisions that a number in this programme changed. If the answer is one or none, the programme is reporting rather than measuring, and cutting it is a capacity win rather than a loss. Then name the removal cost honestly: leadership loses a weekly per-team table it has come to expect, so offer the estate-level trend plus one qualitative paragraph in its place rather than simply taking it away.
When this is the wrong answer
If the platform is six months old with four capabilities and a handful of consumers, this is all far too much machinery. Instrument the funnel for the one capability you are iterating on, talk to the eight teams directly, and leave it there. Telemetry this heavy earns its cost once you can no longer fit your consumers in a room.