CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
5 to work through
-
intermediate
You are sizing the Kafka topic behind a CDC pipeline. The source Postgres takes about 12 million row changes a day across 40 tables, the average row is about 1.2 KB, the connector emits before and after images plus metadata, and the platform team wants 7 days of retention so a failed consumer can be rebuilt without a new snapshot. Roughly how much broker storage, and which assumption dominates the error?
3 min answer -
advanced
A CDC pipeline feeding your warehouse falls three hours behind during a source system's batch job, and the source's transaction log retention is 24 hours. What is the risk and what do you change?
2 min answer -
advanced
A content platform's search index lags the source by up to 30 minutes. Product wants it under 10 seconds. What do you assess?
2 min answer -
advanced
A logistics platform feeds its warehouse from operational databases by change data capture. What must the pipeline design handle?
2 min answer -
advanced
A platform builds change data capture from its operational databases. What is the failure that turns an analytics problem into a production incident?
2 min answer
4 terms in this topic
CDC Pipeline Design
The concerns that turn log-based change capture from a demo into something that can be relied on — snapshots, ordering, schema change and replay.
metricChange Volume Ratio
The proportion of a change stream that is updates and deletes rather than inserts, which decides target storage design, merge cost and whether a tabl…
patternLog-Based Ingestion
Building the pipeline around a database's own change log — an initial snapshot followed by a continuous delta stream, with the two stitched together.
patternSnapshot-Stream Convergence
Starting the change stream before taking the snapshot, so that changes occurring during the backfill are captured rather than silently lost.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.