Telemetry is 28% of infrastructure spend on a contract billed per host and per custom metric, with 14 months left on a minimum annual commitment. The team wants to move to a self-hosted stack. Sequence the move under live traffic, and say when the saving actually starts.
Show the full answer Hide the answer
The sequence
- Put a vendor-neutral collector in the path first, still exporting to the incumbent. Nothing changes operationally, and every later step becomes a routing change instead of a re-instrumentation project. This is the only step whose order is not negotiable, and it is reversible by reverting one configuration.
- Measure what is used before moving anything. Query the vendor's own metadata for which custom metrics back a dashboard or an alert. In most estates a minority do. Delete the rest now, because migrating unused series makes you pay for the waste twice during the overlap, and it may cut the bill enough to change the decision.
- Dual-run one non-critical service end to end, comparing alert for alert rather than dashboard for dashboard. Dashboards are opinions; alerts are the contract.
- Move by signal, not by service. Metrics first, because they are the smallest volume and carry most alerts. Logs next, because they are the largest volume and the easiest to tier. Traces last, because they are hardest to reproduce and most valuable mid-incident.
- Cut alert routing over only after the new stack has carried a real incident, keeping the old alert set live until it has.
- Decommission at the contract boundary, not when the engineering work finishes.
Where data can diverge, and how you would know
The same metric in two systems will disagree whenever aggregation windows, retention or cardinality handling differ, and the disagreement is invisible until an incident. Run a set of golden queries daily against both systems for three services (error rate, p99 latency, saturation) with an agreed tolerance, and treat a breach as a migration blocker. Counters reset differently across collector restarts, which is the most common source of a quiet mismatch.
The point of no return
Removing the vendor agent from production hosts. Everything before that is a rollback by configuration. Schedule it after the new stack has survived an incident and a deploy of the observability stack itself, because a telemetry system that cannot be upgraded safely will be the cause of your next blind spot.
When the saving actually starts
Not when the migration finishes, but at the contract boundary, and only above the minimum commitment. Under a minimum annual commitment, reducing usage saves nothing at all until renewal. So the correct pace is to land the migration one or two months before renewal, not as fast as possible. Meanwhile the run rate goes up: you pay the vendor floor plus the new stack's storage, ingest and the on-call attached to it. Budget a dual-run period of roughly a quarter and say so in the business case, because a cost programme whose costs rise for two quarters loses support at exactly the wrong moment if nobody warned about it.
When this is the wrong answer
If the bill is driven by cardinality you have never controlled, you will rebuild the same bill in your own storage and add an on-call rota to it. Fix cardinality and retention first, then re-price. Teams that do this often find the vendor bill falls far enough that the migration stops paying, which is the cheapest possible outcome. As a rough threshold, running your own telemetry platform starts to make sense somewhere above a few hundred thousand a year of vendor spend; below that the engineering time dominates every other term in the model.