Change Data Capture Pipeline  ·  View 18 of 21  ·  Operations

Sink Lifecycle

Register, backfill, serve, verify, rebuild, swap, retire — a loop that closes because rebuild is routine.

Editable source SVG draw.io All views
Register shape and freshness Backfill chunked snapshot Serve inside freshness target Verify daily reconciliation Rebuild to shadow log or archive Swap in atomic, no partial reads Retire no consumer, no cost Sink lifecycle offset parked replay over snapshot as-of published divergence or bad transform diff against live offsets realigned registry entry removed Sink Lifecycle — Rebuild as a Routine, Not a Disaster Application we own Data store Security / platform Decision point Opportunity The loop closes because a rebuild is the same code path as a backfill, run against a shadow table. v 1.0 · owner Data Platform Architecture · date 2026-10

Why a loop

  • A sink is never repaired in place. Divergence or a bad transform sends it back to rebuild, and rebuild is the same code path as the original backfill (ADR-11).
  • The swap is atomic into a shadow table, so a rebuild never serves partial state.
  • Retirement is on the loop deliberately: a sink nobody reads is a cost with a tenant-data footprint.

Assumptions

  • Every tier-1 table's largest sink is rebuilt from the archive quarterly and diffed against live, so the 8-hour rebuild figure is measured rather than hoped.
  • Verification is daily for tier-1, weekly for tier-2.

Risks

  • A rebuild competes with live apply for sink write capacity; the MVP throttles it rather than scheduling it, which is the cruder of the two options.