metric

Backhaul Cost per Device-Day

also called Bytes per Device-Day, Per-Device Data Budget

The bytes a single device sends and receives in a day multiplied by the price of the link - the figure that converts a sampling decision into a line in the connectivity contract.

costtelemetrycellularsamplingcapacity

A firmware engineer adds one diagnostic field at a one-second cadence. On the bench it is 40 bytes, invisible. Across 200,000 devices on cellular it is a renegotiation of the connectivity contract and a conversation with a finance team who were not told about the field.

Bytes per device-day is the unit that makes those two views the same conversation. It is deliberately per device, because that is the number a plan is priced in and the number an engineer can reason about, and it scales to the fleet by multiplication.

Why it matters

Telemetry volume decisions are made by whoever writes the sampling code, usually without a price in front of them, and they are extremely hard to reverse: once dashboards, alerts and models consume a stream, reducing its rate is a negotiation with every consumer.

The metric also decides architecture, not just cost. It is the first input to where processing happens. If raw data cannot affordably leave the device, filtering and aggregation move onto the device or a local gateway, which changes the entire shape of the system. Deciding that after the ingest pipeline is built is the expensive order.

Implementation patterns

  • Count the wire, not the payload. Headers, TLS records, acknowledgements and keepalives commonly exceed the measurement itself. A 400-byte update every 5 seconds over an 8-hour shift is about 2.3 MB per device-day, near 70 MB per device-month.
  • Include downlink. Commands, configuration pulls and firmware are often the larger half. An 8 MB image once a month is 0.27 MB per device-day on its own.
  • Budget it per stream and publish the budget, so adding a field means spending from an allocation rather than appending to a struct.
  • Put the number in the pull request template for firmware that changes sampling, with the fleet multiplication already applied.
  • Alert on the metric per device cohort, not in aggregate. A regression usually appears on one firmware version first, where a fleet-wide average hides it for weeks.
  • Convert to currency with the real tariff, including per-session or per-message charges on NB-IoT plans, which can dominate byte charges for small frequent messages.

Industry example

Mobility and logistics platforms operating across markets with patchy coverage — the problem a dispatch platform of Grab's kind faces by construction — pay for connectivity per device per month in bulk contracts, so the design question is never "can we store this" but "does this fit the plan". The pattern repeats in consumer IoT: a vendor shipping a sensor with a lifetime cellular subscription bundled into the purchase price has fixed its revenue per device at manufacture, and bytes per device-day is then a direct subtraction from product margin for the next five years.

Failure scenarios

  • A sampling rate raised in a firmware release pushes a fleet past its pooled plan mid-month, and the overage is charged per megabyte at retail rates.
  • A retry loop on a protocol error multiplies traffic by ten for the devices in the worst coverage, which are also the ones with the most expensive roaming.
  • Keepalives for a feature nobody uses. A 60-second keepalive at 100 bytes each way is 4 to 5 MB per device-month carrying no information.
  • Debug logging left enabled for a cohort after an investigation, found when the bill arrives.
  • A plan priced on average usage meeting a distribution with a long tail: the devices in poor coverage generate several times the mean, and the mean was what was bought.

Trade-offs

Spending bytes buys answers. A fleet that samples aggressively can diagnose remotely, which raises the share of issues resolved without a site visit, and a visit costs on the order of $150 to $300 in 2025. Cutting data to the bone saves a connectivity bill and can raise a much larger support bill, so the metric is only meaningful alongside remote-resolution rate.

The honest formulation: spend bytes on the data that changes a decision, and charge the rest to the device. On-device summarisation, event-triggered detail and store-until-asked diagnostics all reduce the daily number without reducing what can be known when something goes wrong.

When not to use it

On mains power and customer Wi-Fi, bytes per device-day stops binding and ingest, storage and query cost take over, which is a different calculation with different winners — there, cardinality and retention decide. It is also the wrong metric for a fleet of a few hundred devices, where the entire connectivity bill is smaller than one engineer-week and the effort belongs elsewhere.

Interview question

Q: A product manager wants a new 50-byte field sampled every second from 200,000 devices on cellular. Give me the number, and tell me what you would propose instead.

What a strong answer covers: the wire arithmetic including framing (about 4 MB per device-day, roughly 120 MB per device-month, around 24 TB a month fleet-wide), what that does to a plan priced at tens of megabytes per device, and alternatives: event-triggered capture at full rate with summaries otherwise, on-device retention with upload on request, and a cohort rollout to measure before committing.

Quick check

Quiz: Why is bytes per device-day computed on the wire rather than from the payload? Answer: because headers, TLS records, acknowledgements, retries and keepalives routinely exceed the measurement, and plans are billed on the wire.

Flashcard: Which number connects a sampling decision to the connectivity contract? — Bytes per device-day multiplied by fleet size and tariff; count downlink and keepalives, and alert per firmware cohort rather than in aggregate.