advanced 2 min answer

A field-service app syncs with a change cursor. Support reports that three customers see cancelled jobs reappearing months after they were closed, and every affected engineer had been on extended leave. The sync endpoint returns 200 for these devices and the client reports no errors. Where do you look, in what order?

synctombstonescursorsretentiondata-integrity
Show the full answer Hide the answer

The first three things to look at

  1. The retention policy on the change feed, next to the distribution of cursor ages the server is accepting. Change tables are pruned to stay fast, commonly at 30 days. Plot p99 and maximum accepted cursor age against that number.
  2. Whether the server validates cursor age at all. A cursor older than the pruning horizon cannot be answered correctly, and the only safe response is "resynchronise from scratch".
  3. The upload path's treatment of a row the server has never seen. If an unknown primary key is handled as a create, the client does not merely keep stale rows; it reinserts them for everyone.

The diagnosis

Deletions are expressed as tombstones in the change feed, and the feed is pruned after 30 days while a device can be absent for eight months. The returning device presents a cursor from before the prune, the server honestly reports every change since that point, and the deletions are no longer in that range. The device therefore never learns the jobs were cancelled, and on its next push it offers them back as live records. Nothing errors because nothing is broken from either side's point of view: the client asked a reasonable question and got a complete answer to a different one.

The misleading signal

Sync success rate and HTTP status. Both are perfect, which steers the investigation towards the client. The correlation with extended leave is the real clue, and it is in the HR system rather than in any telemetry the team owns.

The fix

  • Make the horizon explicit and enforce it. The server stores a watermark for the oldest change it can still describe and replies resync_required to any older cursor. The client discards local state for that scope and takes a full snapshot. This is the change that makes the bug impossible rather than unlikely.
  • Keep tombstones longer than the longest plausible absence. Measure it: the p99.9 of device reconnect gaps, which for seasonal or leave-heavy workforces runs to six or nine months. Tombstones are cheap, on the order of 40 bytes each, so 2 million deletions is about 80 MB. The pruning job was saving storage that was never worth saving.
  • Never treat an unknown key on upload as a create. A client-generated creation carries a creation marker; everything else is an update that can legitimately fail.
  • The alert: a histogram of accepted cursor ages, firing when any accepted cursor passes 80% of the retention horizon. It fires weeks before a customer sees anything.

When this is the wrong answer

If every client syncs at least weekly and the horizon is a year, skip the resync protocol and keep tombstones forever. The protocol earns its complexity only when the gap distribution has a long tail you cannot shorten, which is the normal case for shared handsets, seasonal staff and anything issued per vehicle rather than per person.