Change Data Capture Pipeline  ·  View 11 of 21  ·  Data

Storage Zones

Four zones, ordered by what happens if the data is lost — only two have a durability claim of their own.

Editable source SVG draw.io All views
Zones ordered by what happens if the data is lost Not ours — operational truth, read-only to this platform PostgreSQL primaries application-owned WAL retention the hard deadline System of record for change — losing this loses history Change log 7 days · replayable Changelog archive 13 months · tiered Control state — small, exact, must survive Stream and sink registry Spanner Schema versions effective LSN Sink offsets Snapshot chunk state Operator audit log 13 months Derived — rebuildable from the log, no retention claim Current-state mirror BigQuery Changelog tables BigQuery Search index Read caches re-publish past 7 days replay Storage Zones — Ownership and Rebuildability External / third party Risk / gap Queue / topic Data store batch event / async Only two zones have a durability requirement of their own; the bottom zone is deliberately disposable. v 1.0 · owner Data Platform Architecture · date 2026-10

Decisions

  • The change log plus the archive are the system of record for change; the mirror, the changelog tables, the index and the caches are disposable (ADR-01).
  • Control state is small, exact and must survive: offsets, chunk progress, schema versions and the audit log.
  • Source WAL retention is drawn as a risk, not a store: it is a deadline this platform does not control (ADR-03).

Assumptions

  • Log 7 days; archive 13 months, tiered to colder storage after 30 days.
  • A full rebuild of the largest table from the archive completes within 8 hours.

Consequence

  • A consumer that falls behind retention needs a re-snapshot, not a retry. That is the price of the seam, and it is an operational deadline rather than a bug.