advanced 3 min answer

A catalogue search index is rebuilt every night by a full extract of a 60-million-row product table. You must move it onto a change-data-capture stream with no search downtime and no window where the index is silently wrong. Sequence it so each step is reversible, and say where the point of no return is.

cdcsearch-indexingdivergencereversibilityzero-downtime
Show the full answer Hide the answer

The sequence

  1. Stand up the connector with its snapshot going to a new topic, consumed by nobody. Before anything else, alarm on the source database's replication slot size. The database retains write-ahead log that the connector has not confirmed, so a stalled consumer becomes a disk-full failure on the primary. Set the alarm at a level that still leaves hours, not minutes.
  2. Build a shadow index from snapshot plus stream. Same mapping, same analysers, separate alias. It serves no traffic.
  3. Measure divergence continuously. Two instruments, because they catch different things: a nightly full-scan checksum per shard over (id, version, updated_at), and a sampled field-level compare of a few thousand random documents every 15 minutes. Declare the budget before you start, for example fewer than one mismatch in 100,000 documents for three consecutive nights.
  4. Mirror queries. Send a copy of production queries to both indexes and compare the top 10 result ids. Checksums catch missing documents; only query comparison catches a document that is present and ranks differently because a field arrived in a different shape.
  5. Flip reads behind a flag at 1%, then 50%, then 100%. Each step reverses in seconds because the old index is still being rebuilt nightly.
  6. Keep the nightly extract running for two more weeks as a repair tool rather than as the source of truth. It is the cheapest possible rollback and it costs one batch window.

Where data can diverge, and how you would know

  • DDL that produces no row events. A column added with a server-side default can materialise values in the table without emitting per-row changes on some connectors. The nightly checksum catches it; the 15-minute sample probably does not.
  • Delete representation. The connector emits a delete as a key with a null value, and an indexer that ignores nulls leaves the document in place forever. This is the single most common silent divergence.
  • Transaction boundaries. Row-level change events arrive independently, so the index can briefly hold a state no reader of the database ever saw, such as an order line without its order.
  • The snapshot-to-stream seam. Records changed during the snapshot must be re-applied from the stream, which requires the snapshot to be read at a known log position and the stream consumed from that position.

The point of no return

Not the read flip. It is deleting the nightly extract job, and the moment the source's log retention falls below the time needed to rebuild from the stream alone. Up to that point rollback is a flag. After it, rollback means rebuilding a pipeline under incident pressure. Make that deletion an explicit decision with a date, not something that happens because a cluster was reclaimed.

When this is the wrong migration

If the nightly rebuild completes in 40 minutes and the business tolerates day-old catalogue data, keep it. The full rebuild is self-healing by construction: every bug is corrected within one night, and no divergence can persist. CDC buys minutes of freshness and costs a permanent coupling to the source table's physical schema, which means the catalogue team can no longer rename a column without breaking search. Stream it when a price or stock change has to be searchable in under a minute, and say who loses money when it is not.