Change Data Capture Pipeline  ·  View 07 of 21  ·  Structure

Layered Architecture

Seven layers, one rule: only Capture may read the source, and it only ever reads.

Editable source SVG draw.io All views
Consumers Dashboards Product search Fraud service Partner webhooks Serving Current-state mirror BigQuery Changelog tables append-only Search index Invalidation topic Projection Warehouse applier micro-batch Search indexer Fan-out worker Stateless transform per event Transport Change log 7-day · key-hash Dead-letter topic Per-sink offsets Capture Log reader logical decoding Snapshot chunker Relation detector schema drift Column masker Sources PostgreSQL primary WAL Read replica Heartbeat table 10 s Control and ops Stream and sink registry Schema registry Lag and reconciliation Operator API and audit WAL chunked read change events pull by offset upsert SQL Layered Architecture External / third party Data store Queue / topic Application we own Interface / broker Security / platform synchronous batch event / async The only layer that may read the source is Capture, and it only ever reads. v 1.0 · owner Data Platform Architecture · date 2026-10

Why layered at all

  • The layering is the dependency rule made visible: nothing above Transport may hold a connection to a source database.
  • Control and operations sits beside every layer rather than on top, because pausing a table is an action on Capture and resetting an offset is an action on Transport.

Assumptions

  • Transform is stateless and per-event by construction; joins and aggregation belong to the sink (ADR-09).
  • Offsets live with the log, not with the applier, so an applier is replaceable without coordination.

Omitted

  • Fan-out and search edges above the Projection layer, to keep the flow one-way.
  • The archive, which is a Transport-layer store shown on views 10 and 11.