metric

Projection Rebuild Budget

also called Read Model Rebuild Time, Rebuild Service Level Objective

The measured wall-clock time to rebuild a read model from scratch, treated as a service level objective, because it bounds how quickly a projection defect can be corrected and therefore whether CQRS is reversible in place.

cqrsprojectionread modelrecoverycapacity

A projection bug ships on a Monday and is found nine days later. The write model is correct, the log is intact, and the read model has been answering wrongly for 216 hours. Someone asks the only question that matters operationally: how long to rebuild it?

If nobody has measured, the answer is a guess and the guess decides the incident. A team that believes the rebuild is two hours will start it and discover at hour nine that it is not finished and that search has been degraded all day.

The rebuild budget is that number, maintained as a property of the system rather than discovered under pressure. It is measured, and re-measured as history grows.

Why it matters

Rebuild time is the recovery time objective of the read side, and most teams adopting CQRS never write it down. Everything else follows: whether a read-model schema change is routine or a project, whether a projection defect costs an afternoon or a week, whether a second read model is cheap to add.

The number also grows silently. A projection that rebuilt in forty minutes at launch rebuilds in eleven hours three years later, and nothing alerts because nothing measures it between incidents.

Implementation patterns

  • Measure it on a schedule, not on demand. A monthly rebuild into a throwaway index gives the real number and exercises a code path otherwise only run in an emergency.
  • Know which resource binds. The log's read throughput is almost never the constraint; the read store's write capacity is. A single worker doing one indexed upsert per event sustains on the order of 2,000 to 10,000 events a second, and a batched search or columnar sink roughly an order of magnitude more.
  • Measure the event-to-write fan-out, because it divides every estimate. One event touching a document, an aggregate and two counters costs four writes.
  • Cap parallelism at the partition count, so partition count is a rebuild-time decision as much as a throughput one.
  • Snapshot periodically so a rebuild starts from a checkpoint rather than the first event, and rebuild into a parallel index with a pointer flip once the budget exceeds tolerable degradation.

Industry example

The New York Times publishing pipeline, described on the Confluent engineering blog in 2017, is the clearest public case of this constraint being designed for rather than discovered. The whole published archive back to 1851 was under 100 GB, which is why replaying the log from the beginning to rebuild a downstream store was a normal operation. The design works because the budget is small; the same shape over a clickstream has a budget measured in days and the capability becomes theoretical.

Failure scenarios

  • The number is never measured, so the first rebuild is attempted during an incident and its duration learned the hard way.
  • The rebuild saturates the read store and takes live queries with it, converting a correctness problem into an availability one.
  • No targeted replay path exists, so a bug affecting 400,000 keys out of 90 million forces a full rebuild because it is the only operation available.
  • Rebuild is not idempotent, so an interruption at hour six leaves a partial index worse than the wrong one it replaced.
  • Reconciliation is missing entirely, the deeper failure above: nine days of drift means nothing compared the write model with the read model.

Trade-offs

Keeping the budget small costs money and complexity - snapshots, a second index, a deployment path that addresses two index versions. Letting it grow costs nothing until the day it costs a multi-day degradation.

Choose Gains Pays
Rebuild in place No duplicate storage or routing The read side is degraded for the whole budget
Parallel index and pointer flip Rebuild becomes a background job Double read-store storage and a version-aware read path

When not to use it

If the read model is a database materialised view the engine refreshes, the budget belongs to the engine and measuring it yourself is theatre. The same applies when the read model rebuilds in minutes.

If the only reason for CQRS was that reads and writes had different shapes, a replica plus an index may serve the same purpose with no projection to rebuild. The cheapest rebuild budget is the one you do not have.

Interview question

Q: Your service has a read model built from 1.2 billion events across 24 partitions. Give me a range for the rebuild, and the one measurement you would take before quoting it to a VP.

What a strong answer covers: that read-store write capacity binds rather than log throughput; a range rather than a point, from roughly 11 hours at 30,000 writes a second to roughly 67 hours single-threaded; the event-to-write fan-out as the measurement dominating the error; and the decision it drives, which is in-place rebuild versus a parallel projection.

Quick check

Quiz: Which resource usually binds a projection rebuild, and what does that imply about parallelism? Answer: The read store's write capacity, not the log's read rate - so parallelism is capped by partition count and by how much write capacity you take from live queries.

Flashcard: Why is rebuild time the number that decides whether CQRS is reversible? - Because if a rebuild takes longer than the business will tolerate a wrong read model, you cannot fix a projection in place and must run a second live projection instead.