# Change Data Capture Pipeline

**Solution Architecture v1.0 · Google Cloud · Data Platform Architecture · 2026-10 · 21 views · 15 architecture decision records**

A customer cancels a subscription. That edit lands in exactly one place — an operational PostgreSQL database — and the product shows it in five: the analytics dashboard the board reads, the search box, the "your plan changed" email, fourteen partner webhooks, and the support console an agent is looking at while the customer is still on the phone. Something has to carry the change from the one place it was written to the many places it is read. When that something is a nightly batch job, users call the product broken; when it is a replication job written per consumer, the fifth consumer costs what the first did all over again and the first bad transformation is repaired with hand-written SQL against production. This package is that carrier, built for an assumed mid-market SaaS: 4,000 tenants, 12 operational PostgreSQL databases, 180 replicated tables, 9,000 row changes a second with a 4× burst for 30 minutes, 780 million change events a day, and a p95 commit-to-sink lag of five seconds — on Datastream for log-based capture, Pub/Sub as the change log, Dataflow for transform and apply, BigQuery as the mirror and changelog sink, Cloud Storage for a thirteen-month archive, and Spanner for control state.

The product is not the ability to copy rows quickly. It is **the ability to be wrong and recover without a human writing SQL** — because a derived store will eventually disagree with its source, and the only question that matters is whether that is a rebuild or an incident.

The design rests on one rule: **the change log is the system of record for change, every sink is a disposable projection of it, and capture never knows its consumers.**

The decisions that carry the design:

- **The log is the record; sinks are projections.** The mirror, the changelog tables, the index and the caches are all derived and rebuildable, so a bad transform, a corrupt partition, a wrong coercion and an operator error have one remedy instead of four disaster procedures — and the fifth sink costs an offset, not a change to capture (ADR-01).
- **Read the write-ahead log, never poll the table.** Polling cannot see deletes, misses intermediate versions and invents its own ordering; application dual writes fail silently in the one case that matters, a crash between the commit and the publish (ADR-02).
- **The replication slot is shared fate with the source.** An abandoned slot fills someone else's disk, so slot headroom is a first-class alarm, the platform refuses a pause that would outlive retention, and a position gap stops the table loudly rather than resuming across it (ADR-03).
- **Exactly-once lives in the data model.** At-least-once delivery, idempotent apply keyed on (primary key, log position), a monotonic per-key sequence, and the offset committed after the sink write — which turns a crash into a duplicate rather than a gap (ADR-04).
- **Chunked snapshot, stream started first.** Record the position, stream from it, snapshot in primary-key ranges interleaved with live traffic, and let monotonic apply resolve the overlap. No lock, no long transaction, resumable by chunk, and one apply path for backfill and live capture (ADR-05).
- **The source's write path is sacred.** Snapshot reads come from a replica inside a declared source budget, and the backfill is the thing that slows down when the source is busy — never the application (ADR-06).
- **Mask before the log, not after.** A regulated value that reaches the change log lives in the log, the archive and every sink built from it; capture-side exclusion is the only placement where it never arrives (ADR-07).
- **Schema change is found in the stream.** Compatible changes propagate within a minute with no human; breaking ones pause exactly one table within ten seconds. An unregistered column is propagated or excluded — never silently dropped, which is the failure that looks correct (ADR-08).
- **Stateless transform, clever sink.** Projection, coercion, masking and tenant stamping per event; joins and aggregation in the sink. The pipeline holds no business state, so it is restartable from any position with nothing to recover but an offset (ADR-09).
- **Deletes are tombstones by default.** Hard delete matches the source and destroys last Tuesday; the reversible representation is the default and the other is a per-table opt-in (ADR-10).
- **Rebuild, never repair.** A divergent sink is replayed into a shadow table and swapped atomically, and rebuild is rehearsed quarterly — which is what makes the eight-hour figure a measurement rather than a hope (ADR-11).
- **Divergence is found before a consumer finds it.** Row counts and sampled checksums per tier-1 table daily, with reconciliation coverage reported as a number, because the cheapest improvement available to this design is coverage rather than faster capture (ADR-12).
- **Lag is measured on a heartbeat, end to end, per sink.** Time-since-last-event reports zero for an idle table and for a stopped stream, which are opposite conditions; a heartbeat row every ten seconds tells them apart (ADR-13).
- **Control state never lives inside a projection.** Offsets, schema versions, chunk progress and audit are a few million exact rows and belong apart from the billions of disposable ones — you do not store the instructions for rebuilding something inside the thing you may have to rebuild (ADR-14).
- **One capture region, stateless standby.** One source log has one position sequence, so active-active capture is a disagreement rather than redundancy; and a position after a source failover may not mean what it meant, which is reconciled explicitly rather than derived from a timestamp (ADR-15).

The architecture one-pager (including why the design should still hold up in ten years, and the five risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~13 min) and [docs/decision-record.md](docs/decision-record.md) (~50 min).

---

## What is here

| Path | Contents |
|---|---|
| `diagrams/index.html` | The landing page: 21 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 15 decision records, the capability-to-technology table and the package glossary as markdown |
| `specs/part-a..c.json` | Diagram specifications, the source of truth for every view |
| `specs/manifest-a..b.json` | Acts, page titles, subtitles and reasoning cards |
| `specs/adr-onepager.json`, `specs/adr-records-a..b.json` | The one-pager, the decision records, the capability-to-technology table and the glossary |
| `scripts/build.sh` | Rebuilds every deliverable from the specs (Node only, no network) |
| `ask.md` | The requirement |

## The twenty-one views

| # | View | What it answers |
|---|---|---|
| 01 | System Context | One writer, many readers — and the three dependencies that can stop the pipeline, one of which is somebody else's disk |
| 02 | High-Level Architecture | Five stages and one seam: everything left of the log reads the source, everything right of it reads the log |
| 03 | Actors and Their Journeys | Four human questions and two machine ones, including the engineer who can break this platform by doing their job correctly |
| 04 | Journey — Wire Up a New Sink | The trough is not registration: it is pressing go on a 400 GB backfill against a live primary with no ETA |
| 05 | Journey — Ship a Schema Migration | Nobody tells the pipeline a migration is coming, so the design assumes it finds out from the stream |
| 06 | Journey — Chase a Wrong Number | Stale, wrong or right in under a minute — and the coverage gap that makes the question hard |
| 07 | Layered Architecture | Seven layers and one rule: only Capture may read the source, and it only ever reads |
| 08 | Platform Components | Four planes, and why the control plane is a separate store from everything it describes |
| 09 | Integration Surface | Three interfaces in, three out, and exactly one contract in each direction |
| 10 | Data Flow — One Row Change | The three fields everything downstream relies on, and the two decisions every applier makes before it writes |
| 11 | Storage Zones | Four zones ordered by what happens if the data is lost; only two have a durability claim of their own |
| 12 | Control-Plane Data Model | Ten entities, and the composite key that makes a replayed event a no-op |
| 13 | Commit to Sink in Five Seconds | Thirteen messages, and the one ordering rule that turns a crash into a duplicate rather than a gap |
| 14 | Snapshot to Stream Handoff | The initial-snapshot problem solved with no lock and no gap, and what happens when it is interrupted |
| 15 | Schema Drift — Four Paths | Five change classes, two of which stop exactly one table, and the one outcome that is forbidden |
| 16 | Deployment Architecture | One write region, a standby that holds no state, and the two stores that are deliberately multi-region |
| 17 | Observability | Six signal families across five stages, reduced to four alarms that map to four distinct actions |
| 18 | Sink Lifecycle | Register, backfill, serve, verify, rebuild, swap, retire — a loop that closes because rebuild is routine |
| 19 | Security Trust Zones | Four zones, and the one control that has to sit upstream of the log to work at all |
| 20 | Identity and Access Flow | A privileged action the platform is allowed to refuse, and why the refusal is the control |
| 21 | Failure Classes | Nine classes with their blast radius, and not one recovery that is hand-written SQL |

## Rebuilding

```bash
bash scripts/build.sh
```

Node 20+ and nothing else. No draw.io Desktop, no browser, no network. The build assembles
`specs/views.json` and `specs/manifest.json` from the authoring parts, generates the draw.io
sources, fails on any geometry or routing defect, renders the editable and plain SVGs, writes the
HTML pages and the index, injects the one-pager and the decision record, and proves every relative
link in `diagrams/` resolves.

At v1.0 the geometry gate reports 0 errors and 0 warnings across 21 files, the routing gate
0 errors and 2 clutter warnings across 153 edges, all 344 nodes resolve to an icon with no weak
matches, and the link check passes on 166 relative links. The two warnings are label stacking on
the integration and trust-zone views, where several relationship edges converge on one node; they
were reduced by shortening labels, merging two sink nodes and dropping edges other views already
carry, and removing the remainder would mean dropping a source interface or a zone-crossing edge
that is worth more to the reader than the warning costs.

## A note on the numbers

Every rate, latency, volume, retention and cost figure in this package is a **stated assumption**,
chosen to be defensible and arguable rather than measured. They are stated precisely so that a
reviewer can disagree with one and follow it to the decision that depends on it. `ask.md` marks
them as assumptions section by section; the decision record's evidence note restates the operating
context in one paragraph.
