advanced 4 min answer

A platform team runs a vendor-specific agent across 2,000 services feeding a metrics store holding billions of active series — the scale eBay described for Sherlock.io when it published its move to OpenTelemetry in 2022. They want to migrate to OpenTelemetry without a telemetry gap and without asking 400 application teams to coordinate. Give the sequence, and say where the data can silently diverge.

ebayopentelemetrymigrationdual emitsemantic conventionscollector
Show the full answer Hide the answer

The sequence

Each step is reversible on its own, which is the property that makes a migration of this size survivable.

  1. Stand up the collector tier first, receiving nothing. The OpenTelemetry Collector speaks OTLP and the old protocols, so it can sit in the path before any application changes. Size it, alert it, run it empty for a week. A migration that begins by changing application code has already lost the ability to roll back cheaply.
  2. Point the existing agents at the collector, passing through unchanged. Every byte now flows through the new tier to the existing backend. No semantics have changed, and you have validated the hardest operational piece — capacity, backpressure, failure behaviour — while the output is provably identical.
  3. Fan out at the collector, not in the application. Add a second exporter to the new backend. Both now receive the same data from the same source. Doing the dual emit in the collector rather than in 2,000 services is what removes the need for 400 teams to coordinate.
  4. Reconcile, with the old backend as the reference. Run both for at least one full weekly cycle, comparing per-metric rather than in aggregate.
  5. Move the alerts before the dashboards. Alerts are the thing whose silence is dangerous. Leave the old ones running in parallel to a low-priority channel for a fortnight; any discrepancy is a migration bug found by machine.
  6. Migrate dashboards and saved queries. The long, dull tail, and where the schedule actually goes: a few thousand saved queries at a few minutes each is weeks, and it parallelises badly because each needs its owner to confirm it still means what they thought.
  7. Replace the agent with the OTel SDK service by service, as teams touch their code anyway. The collector normalises both, so this has no deadline and no coordination. Teams that never migrate keep working, which is what makes it politically survivable.
  8. Only then decommission. The point of no return is turning off the old ingest, months after the last alert moved.

Where the data can diverge, and how you would know

This is the part that is underestimated. Five specific divergences, in rough order of how often they bite:

  • Histogram bucket boundaries. The two will not choose the same defaults. Every percentile from a histogram is an interpolation within a bucket, so different buckets give different p99s from identical traffic, often by 10–30% at the tail. This looks like a latency regression on cutover day and is not one. Pin the boundaries explicitly on both sides before comparing anything.
  • Attribute naming. Semantic conventions rename nearly everything: http.status_code versus status, service.name versus a vendor tag. Every dashboard, alert and recording rule on the old name silently matches zero series. A query matching nothing returns "no data", and most alerting treats no data as not-firing, so this divergence produces an alert that never fires again and is noticed during the incident it was meant to catch.
  • Temporality. Cumulative versus delta counters. Feed a delta counter to a backend expecting cumulative and the rates are nonsense in a way that looks plausible on a graph.
  • Unit changes. Seconds versus milliseconds is the classic. A dashboard showing 0.045 where it showed 45 is caught instantly; an alert threshold of > 200 now comparing against seconds is not caught at all.
  • Dropped data under load. The collector has queues and drops under a burst where the old path may not, so the backends disagree during exactly the periods you care about.

How you would know: a reconciliation job, not an eyeball. For the top few hundred metrics, compute the 5-minute aggregate from both backends and alert when they differ by more than a few percent for two intervals. A few days of work, and the difference between finding divergences during the migration and during an incident six months later.

How long it really takes

For a fleet of this size, the collector tier and dual-emit are weeks. The dashboard and alert tail is quarters. The agent replacement across 2,000 services is years and should be explicitly planned as never finishing — which is fine, because after step 3 the collector makes the source irrelevant. Teams that promise a completion date for step 7 are the ones who later force a flag day, and a flag day on telemetry means going blind on purpose.

When not to do this at all

Portability and a vendor-neutral wire format are strategic benefits, not urgent ones. If the existing agent works and no renewal depends on it, this is quarters of platform effort for no change a user can perceive. The case flips on one of three things: a second backend is genuinely wanted, instrumentation is being written from scratch for a new language tier anyway, or the vendor agent blocks a runtime upgrade. Absent one of those, do steps 1 to 3, which cost little and buy the optionality, and stop there deliberately rather than by drift.

The rollback at each stage

Steps 1–3 roll back by removing an exporter. Step 5 rolls back by re-enabling the old alert routing. After step 8 there is no rollback, which is why it is last and why the gap before it should feel excessive.