Change Data Capture Pipeline  ·  View 10 of 21  ·  Data

Data Flow — One Row Change

What a change event carries, and the two decisions every applier makes before it writes.

Editable source SVG draw.io All views
Committed change Row commit INSERT · UPDATE · DELETE WAL record LSN + txid Decode and enrich Decoded change pre + post image Mask or exclude declared columns Stamp metadata tenant · schema v Log Change log partition = hash(PK) Archive writer 13 months Fan-out Warehouse offset Search offset Fan-out offset Materialise Micro-batch 5 s or 10k rows Idempotent apply key = PK + LSN Discard if stale LSN ≤ current Sinks Current-state mirror Changelog table Search index continuous Data Flow — One Row Change, End to End External / third party Application we own Security / platform Queue / topic Data store Decision point batch The post-image, the primary key and the log position are the three fields everything downstream relies on. v 1.0 · owner Data Platform Architecture · date 2026-10

The three fields that matter

  • Post-image, primary key and log position: idempotence, ordering and staleness detection all derive from these.
  • The schema version is stamped at capture so a replayed event is interpretable years later (ADR-08).
  • Tenant id is derived at capture and never inferred at the sink, because an inferred tenant is a cross-tenant leak waiting for a refactor.

Assumptions

  • Micro-batch flushes at 5 seconds or 10,000 rows, whichever comes first.
  • p99 event ≤ 64 KB; events over 1 MB are dead-lettered with a size error.
  • 780 million change events a day, about 1.1 TB uncompressed.

Risks

  • Masking at capture is irreversible by design: a column excluded here cannot be backfilled without a re-snapshot.
  • Discarding a stale event depends on a monotonic sequence per key; a sink that ignores it will overwrite new values with old ones.