Change Data Capture Pipeline  ·  View 13 of 21  ·  Runtime

Critical Flow — Commit to Sink

Thirteen messages, and the one ordering rule that turns a crash into a duplicate rather than a gap.

Editable source SVG draw.io All views
Product service PostgreSQL Capture plane Change log Warehouse applier BigQuery mirror Monitoring 1. UPDATE order SET status 2. commit acknowledged 3. WAL record at LSN 4. decode, mask, stamp schema v 5. publish, key = hash(PK) 6. durable ack 7. pull batch from offset 8. transform, group by key, keep max LSN 9. MERGE micro-batch 10. rows affected 11. commit offset 12. end-to-end lag from commit_ts 13. non-retryable: to dead letter Critical Flow — Commit to Sink in Five Seconds The offset is committed after the sink write, never before: that is what makes a crash a duplicate rather than a gap. v 1.0 · owner Data Platform Architecture · date 2026-10

The ordering rule

  • The offset is committed after the sink write, never before. A crash in between replays the batch, which idempotence absorbs; the reverse loses rows silently (ADR-04).
  • The applier groups a batch by key and keeps the highest log position, so a row changed five times in one batch is written once.
  • Lag is computed from the source commit timestamp carried on the event, which is why the heartbeat exists (ADR-13).

Numbers

  • p50 ≤ 2 s, p95 ≤ 5 s, p99 ≤ 15 s end to end; lag over 60 s for 5 minutes on a tier-1 table pages.
  • Capture adds ≤ 5% source CPU and ≤ 10% to p99 source write latency (assumption, and a budget the snapshot path yields to).

Risks

  • A transform exception on a single event can stall a partition if the retry budget is unbounded; the dead letter is the release valve.
  • The acknowledged commit to the product service is independent of this flow — nothing here can make a user's write slower, which is the whole point.