Change Data Capture Pipeline  ·  View 02 of 21  ·  Context and scope

High-Level Architecture

Five stages, one seam: everything to the left of the log reads the source, everything to the right reads the log.

Editable source SVG draw.io All views
Sources PostgreSQL primary WAL · 1 slot per DB Heartbeat writer every 10 s Capture plane Datastream capture logical decoding Snapshot chunker PK ranges Column masker at capture Change log Pub/Sub change log 7 days · key-hash Dead-letter topic poison events Projection plane Dataflow appliers one job per sink Schema registry Spanner Sinks BigQuery mirror + changelog Search index Cache invalidation Consumers Dashboards Product search Partner webhooks retries spent schema version Change Data Capture Pipeline — High-Level Architecture External / third party Security / platform Interface / broker Application we own Queue / topic Data store failure / alternate synchronous Every sink reads the same log at its own offset; no sink can slow another, or the source. v 1.0 · owner Data Platform Architecture · date 2026-10

The seam

  • Capture publishes; it never knows which sinks exist. Adding the fourth sink is an offset, not a change to the capture path.
  • Back-pressure from a slow sink terminates in the log's 7-day retention and cannot reach the source's write path.
  • The schema registry sits beside the projection plane because a replayed week-old event must be read with the schema that was in force when it was written.

Numbers on this view

  • Log retention 7 days; changelog archive 13 months (assumptions).
  • One replication slot per source database, shared by all 180 tables.
  • End-to-end p95 ≤ 5 s, measured commit timestamp to sink commit timestamp.

Deliberately omitted

  • The control plane and every metrics edge — they appear on views 08 and 17.
  • The archive write path, on view 10.