advanced 3 min answer

An online-grocery marketplace of the kind Instacart operates splits its catalogue service - writes go to the catalogue team's Postgres, and the search-shaped read model is built and owned by the search team from the catalogue's event stream. Read-path availability improves and the catalogue team stops fielding query requests. What has the team given up, and when does the bill arrive?

cqrsread-modelteam-topologyevent-contractcoupling
Show the full answer Hide the answer

What is gained

Real things, worth naming before the bill. The read path scales and fails independently of writes. The query shape no longer constrains the write schema, so the catalogue's normalised model stays normalised. Two teams deploy on their own cadence. None of that is what makes this expensive.

What is paid

The event stream stopped being an implementation detail and became a published interface with a producer and a consumer in different reporting lines. Three bills follow.

Schema changes become release sequencing. Renaming a field is now: producer emits both forms, consumer deploys the reader, producer removes the old form — three deployments across two backlogs. With a two-week train on either side, a change that took a day inside one service takes four to six weeks, and it takes that long for the twentieth such change as well as the first.

The staleness window becomes a cross-team service level objective. When search shows an item in stock that the catalogue retired 90 seconds ago, the shopper arrives at a shelf that is empty. Neither owner's dashboard shows an error: the producer published, the consumer consumed, the projection lag was 90 seconds and nobody had agreed what number was acceptable. The defect surfaces in operations, not in engineering.

Rebuild capability is split from data ownership. The consumer owns the projection and has to rebuild it when its own logic is wrong; the producer owns the retention that makes a rebuild possible. Seven days of retention means a projection bug older than seven days can no longer be repaired from the stream, and the fix becomes a bulk export negotiated with the other team under incident pressure.

When the cost becomes visible

At the first breaking schema change, and at the first incident where the read model is wrong rather than down. Wrongness is the shape to plan for: availability problems page somebody, and semantic drift does not.

How to keep the option to reverse

  • Version the event contract and run the consumer's expectations as tests in the producer's pipeline, so a breaking change fails in CI rather than in the other team's error budget.
  • Publish the staleness window as a number — for example, p99 projection lag under 30 seconds — alerted on by the consumer and visible to the producer, because only the producer can cause some of its breaches.
  • Keep a bulk snapshot path that does not depend on broker retention, and rehearse it. The rebuild time you have measured is the real limit on how fast a projection defect can be undone.

When this is the wrong answer

When both models would be owned by the same team. Then skip the stream: a denormalised read table maintained in the same transaction as the write, or a read replica with a few purpose-built indexes, gives most of the benefit with no contract, no lag SLO and no rebuild procedure. It is the organisational seam that makes CQRS expensive, not the pattern — and the corollary is that the projection boundary is worth placing where you want a team boundary anyway, not wherever the data happens to split.