intermediate 2 min answer

Etsy released StatsD in 2011 alongside a blog post titled "Measure Anything, Measure Everything". The client sends UDP packets and does not care whether anything is listening. What problem forced that design, and what does it have to do with deploying many times a day?

etsystatsdinstrumentationudpdeployment frequency
Show the full answer Hide the answer

The situation they were in

If every engineer is expected to instrument their own code, the cost of adding one metric has to approach zero in both developer effort and request latency. A metrics client that can block, raise, require a schema or need a ticket is a client engineers route around, and uninstrumented code is what makes frequent deployment dangerous: you ship often and learn nothing per ship.

That is the real coupling between measurement and delivery. Deploy frequency is only safe when the feedback per deploy is cheap enough that nobody skips it.

What they chose

A fire-and-forget UDP client and an aggregating daemon. One line of code at the call site, no connection to establish, no acknowledgement to wait for, and no error path for the caller to handle or ignore. The daemon aggregates into buckets and flushes to a time-series backend on a fixed interval.

Sampling is in the protocol, not bolted on. The client can send one packet in ten and stamp the sample rate on the packet; the server scales the count back up before flushing. A developer who wants to count something that happens 50,000 times a second can do so without thinking about the consequences downstream.

Why it fit their constraints

The asymmetry settles it. The failure mode of UDP is losing a measurement; the failure mode of a reliable client is slowing or failing a request. For telemetry, the first is an inconvenience and the second is an outage caused by the instrumentation. Fixed-interval flushing has a second effect that matters at scale: write volume at the backend is bounded by the number of buckets rather than by request rate, so a traffic spike does not become a storage spike.

What it cost them

  • Counts are approximate and the error is silent. A dropped packet looks exactly like an event that did not happen, so you cannot distinguish "no traffic" from "packets lost".
  • Loss rises with load, which is the moment the numbers matter most.
  • There is no per-event detail. Pre-aggregated buckets cannot answer a question about one customer, because the per-request information was discarded at write time.

When not to copy this design

Never invoice, audit or settle from a sampled fire-and-forget counter. Anything a number is billed from needs exact, acknowledged, deduplicated events, and teams that discover this late find it in a revenue reconciliation rather than in a dashboard.

For a small service at modest volume, in-process aggregation with a scraped endpoint is simpler and gives up nothing: no daemon to run, no UDP loss, and exact counts. The design earns its approximations at high request rates and loses to the simpler option below them.

The lesson that generalises

The constraint on measuring everything is usually the friction of adding one metric, not the cost of storing it. The decision rule follows directly: if adding a metric needs a schema change, a review or a ticket, your deploy frequency will be limited by your blindness rather than by your pipeline, and no amount of pipeline work will show up in change failure rate.