advanced 2 min answer

A fleet exports OTLP to a gateway collector deployment with memory_limiter first in the pipeline, then a batch processor, then an exporter with a sending queue. The telemetry backend starts answering in 8 seconds instead of 80 ms, and request volume doubles at the same time. What happens second by second, and what stops it?

opentelemetrycollectorbackpressureload-sheddingqueueing
Show the full answer Hide the answer

Second by second

t+0 to t+10s. Each export takes 100 times longer, so in-flight slots stay occupied. The exporter's sending queue, whose default capacity is 1000 items, fills in seconds at gateway rates. Nothing is visibly wrong yet.

t+10 to t+45s. Batches accumulate in memory behind the full queue and the heap climbs. The memory_limiter crosses its soft limit at limit_mib minus spike_limit_mib and starts refusing data, returning errors to the gRPC callers. This is the component doing its job: it is placed first in the pipeline precisely so it can shed before downstream processors allocate.

t+45s onward. The refusal propagates backwards into every application. A batching span processor sees export failures, its own bounded queue fills, and it drops spans — the correct outcome. The dangerous paths are the ones that do not drop: a simple or synchronous span processor, or a metric reader whose export holds a lock on the instrument registry, adds the collector's failure latency to the application's own request path.

Where it amplifies

Retry policies turn one slow backend into more offered load against it. And the incident becomes harder to see at the moment it matters: gaps appear in every service's metrics, caused by the pipeline rather than by the services, and an on-call engineer reads missing data as a dead service. The failure of the telemetry path is indistinguishable, from the dashboard, from the failure of everything the telemetry describes.

What stops it

Bounded queues everywhere with an explicit drop policy, and retry_on_failure with a maximum elapsed time so the collector abandons work instead of hoarding it. The collector's own metrics — refused and dropped counts per processor, exporter queue size and capacity — shipped somewhere that does not depend on this pipeline. A soft limit low enough that refusal happens well before the kernel's OOM killer, since a killed collector loses everything buffered while a refusing collector loses only the newest arrivals.

Self-healing requires the drop, not the buffer. A pipeline that buffers through a 20-minute backend outage emerges with 20 minutes of backlog and spends another 20 minutes delivering telemetry about the past while new telemetry queues behind it. A pipeline that drops is current the second the backend recovers.

The shape recurs one layer up. OpenAI's December 2024 outage began with a new telemetry service whose Kubernetes API cost scaled with cluster size; per-node DNS caching delayed the visible symptom long enough for the fleet-wide rollout to continue. A rollout gate must outlast every cache and every queue in the path it is gating, or the canary reports on a system that has not yet felt the change.

When this is the wrong answer

For audit, billing or security telemetry, dropping is not available, and the right design is a durable queue in front of the collector with acknowledged writes and an accepted delivery delay measured in minutes. That is a different system with different economics, and mixing it with operational telemetry means paying durability prices for data whose value expires in an hour.