advanced 2 min answer

An inspection app keeps each site's checklist in an observed-remove set CRDT and synchronises by sending the whole replica state on every reconnect. 6,000 devices on cellular, roughly 40,000 elements per site set, about eight reconnects a day. Roughly how many bytes per device-day, and what does the number rule out?

crdtsor-setdelta-statebandwidthmetadata
Show the full answer Hide the answer

The assumptions, stated

  • Payload per element about 60 bytes of actual checklist content.
  • Causal metadata per element: a replica identifier of 16 bytes and a counter of 8, so 24 bytes, plus framing. Call it 100 bytes per live element.
  • Removed elements leave causal residue the replica must keep so an unseen peer cannot resurrect them: assume 30% of live count at about 30 bytes each.
  • Compression of roughly 3 to 1 on the wire, because identifiers repeat.

The arithmetic

Live state is 40,000 × 100 B, about 4 MB. Residue adds 12,000 × 30 B, about 0.4 MB. Roughly 4.4 MB of replica state, near 1.5 MB after compression, eight times a day: about 12 MB per device-day, or 360 MB per device-month.

Across 6,000 devices that is 72 GB a day of sync traffic. The decisive observation is that content is 60 of the 100 bytes, so almost half the bill is metadata, and all of the growth is.

Which assumption dominates the error

Elements per site, then reconnect count. Both are observable today, and both tend to be underestimated at design time because the first year of data looks small.

What the number rules out, and what replaces it

  • A 360 MB per device-month sync budget is incompatible with most pooled cellular plans, which sit in the 500 MB to 1 GB range per device. Sync alone would consume the plan, before telemetry and photographs.
  • Delta-state replication is the direct fix. Ship only the dots a peer has not seen: 40 edits a day at about 100 bytes is roughly 4 KB per device-day, three orders of magnitude less, in exchange for exchanging causal context on every session.
  • That exchange has its own growth. A version vector with one entry per replica is 24 bytes × 6,000 devices, about 144 KB per site, and it grows with the fleet rather than with the data. Keeping it bounded means one sync peer, not many: a star topology where every device syncs only with the server holds vectors at two entries and gives up peer-to-peer sync in the field.
  • Residue needs a retirement policy at this scale, which means a stable causal watermark across replicas, which means knowing which replicas are retired. That is fleet management work, not data-structure work, and it is usually discovered late.

When this is the wrong answer

If one inspector edits a site at a time, which is how field inspection usually works, a per-field last-writer-wins value with a server-assigned timestamp carries zero metadata and is the correct design. A 2019 write-up from a widely used collaborative design tool made the same call in a harder setting: a central server exists, so make it authoritative and skip the coordination-free merge. CRDTs earn their bytes when genuine concurrent editing of the same records is a product requirement rather than a theoretical possibility.