Change Data Capture Pipeline  ·  View 06 of 21  ·  People and journeys

Journey — Chase a Number That Looks Wrong

Stale, wrong or right — and the coverage gap that makes the question hard.

Editable source SVG draw.io All views
Platform SRE with a revenue analyst Goal — Say within minutes whether the dashboard is stale, wrong or right Trigger — "Yesterday's revenue dropped 8% and nothing happened" Done when — A named cause, a repair by replay, and a check that would have caught it 1 · Report a human notices 2 · Triage 3 · Diagnose ◆ moment of truth 4 · Repair ◆ moment of truth 5 · Prevent What they do Analyst raises it Checks lag and as-of Reads reconciliation Compares to source Replays to a shadow Adds the table to tier 1 What the platform shows Freshness per sink Lag within target Checksum divergence Lineage to a transform Replay from position Daily verdict How it feels Calm Tense Exposed Where it hurts Stale or wrong? No check on this table What answers it As-of on every read Four distinct alarms Sampled checksums Per-event lineage Rebuild, not repair SQL Coverage as the output Journey — Chasing a Number That Looks Wrong The trough is a table nobody reconciled: the fix is coverage, not a cleverer query. v 1.0 · owner Data Platform Architecture · date 2026-10

The trough is coverage, not cleverness

  • Triage is fast when lag and as-of are published; it collapses when the table in question has no reconciliation verdict at all.
  • So the architecture treats reconciliation coverage as an output: the set of tier-1 tables with a daily verdict is a reportable number (ADR-12).

What answers it

  • Per-event lineage — source, table, log position, schema version, transform version — explains any single sink row.
  • Repair is a replay to a shadow table and a swap, never correction SQL against the serving mirror (ADR-01, ADR-11).

Assumptions

  • Row-count and sampled-checksum verdict per tier-1 table every 24 hours.
  • Tier-1 is assumed to be 40 of the 180 replicated tables.