advanced 3 min answer

A grocery marketplace shows store availability from a read model fed by change-data-capture streams from each retailer's inventory system. Normal lag is 2 to 4 seconds. During the Saturday peak the stream backs up to 90 seconds and stays there for 20 minutes. Nothing errors and no alert fires. What happens, and which mechanism stops it?

consistencycdcstalenesssubstitutionsfeedback-loop
Show the full answer Hide the answer

Second by second

Shoppers add items that sold out up to 90 seconds earlier. Orders are accepted because acceptance reads the same stale model. The picker in the store reaches the shelf, finds nothing, and starts the substitution flow: contact the customer, offer an alternative, wait for a reply, or refund.

The amplification is physical, not computational. Each stale line item costs a picker 60 to 120 seconds of exception handling. At a peak where pickers are the bottleneck, a few hundred extra exceptions an hour removes real fulfilment capacity, which lengthens delivery windows, which increases cancellations and support contacts, which consumes staff who would otherwise be picking. The software is healthy throughout: queues drain, CPU is unremarkable, error rates are flat.

What the user sees

Two failures with different severities. A shopper who is offered a substitution has a mildly worse experience. A shopper whose order arrives 40 minutes late with three items missing has been given a promise the system could not keep, and the promise was made at checkout, which is the point where staleness stopped being cosmetic.

What stops it

Bound staleness at the decision point, not at the display. Carry the source event timestamp into the read model and check its age where it matters:

  • Browsing and search read the model whatever its age. Availability shown as "in stock at your store" when the data is 90 seconds old is acceptable and always has been.
  • Checkout does not. When the model's age for that store exceeds a threshold of roughly 30 seconds, the reservation call goes to the retailer's authoritative endpoint for the lines in the basket, and if that call fails the item is shown as "availability unconfirmed" rather than promised.

Degrade the freshness of the display and never the correctness of the reservation. That is the whole rule, and it is a per-operation decision rather than a system-wide consistency setting.

What would have to be true for it to self-heal

Catch-up has to be bounded by distinct keys rather than by event count. If the consumer applies every intermediate state for a SKU, a backlog of 400000 events takes as long to drain as it took to create, so lag grows and never recovers within the peak. Compacting by key on the way in — last write wins per store and SKU — makes the drain proportional to the number of changed SKUs, which is how the stream recovers while traffic is still high.

The alert that would have caught it earlier

Not consumer lag in messages, which looks the same at 2 seconds and at 90. Alert on the age of the newest applied event per store at p95, thresholded at the number checkout depends on, and page when a store's data crosses it. Tie the threshold to the business decision so the alert means something specific: past this age, we stop making promises.

When this is the wrong answer

For a marketplace where inventory is warehoused and counts change hourly, an authoritative call at checkout is added latency and a new dependency for no benefit, and the read model alone is correct. The authoritative path earns its cost only where stock turns over faster than the pipeline propagates, which is the defining property of store-level grocery and a handful of other domains.