Change Data Capture Pipeline  ·  View 17 of 21  ·  Operations

Observability

Six signal families across five stages, reduced to four alarms that map to four distinct actions.

Editable source SVG draw.io All views
Source Capture Change log Projection Sink Lag Commit timestamp heartbeat 10 s Capture lag p95 ≤ 1 s Transport lag Apply lag End-to-end lag p95 ≤ 5 s · SLO Throughput Row changes/s 9k steady Decode rate Publish rate Batches/min Rows applied/s Errors Decode failures Dead-letter depth alarm on growth Transform exceptions Apply rejections Saturation Slot retention headroom the disk-full alarm Source CPU added ≤ 5% budget Unacked backlog Worker autoscale Sink quota Correctness Row counts tier-1 daily Snapshot chunk state Position continuity gap detection Schema version in use Sampled checksum divergence alarm Cost Bytes logged per table Worker hours per sink £ per million events target ≤ 0.60 Observability — Signal by Pipeline Stage Four alarms, four actions: lag against target, slot headroom, dead-letter growth, reconciliation divergence. v 1.0 · owner Data Platform Architecture · date 2026-10

The four alarms

  • Lag against the sink's declared freshness target — act on capacity or a stalled job.
  • Replication slot retention headroom — act on the source before its disk fills (ADR-03).
  • Dead-letter growth — act on a poison class or a transform defect.
  • Reconciliation divergence — act on correctness, which is the only one that cannot wait for morning.

Measurement

  • End-to-end lag is measured on a heartbeat row written every 10 s to every source database, so an idle table still reports lag (ADR-13).
  • Cost per million delivered events is a first-class signal with a target of £0.60 (assumption).

Gaps stated

  • There is no source-side error signal: the pipeline cannot see a failed application write, and should not claim to.
  • Capture and log stages carry no cost attribution of their own in the MVP; per-table attribution is Phase 2.