Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
5 to work through
-
intermediate
A logistics platform needs its operational database changes to reach a search index and an analytics warehouse. An engineer proposes writing to all three from the application. What is wrong, and what is the alternative?
2 min answer -
intermediate
A marketplace's search index lags several minutes behind catalogue updates and sellers complain their listings appear stale. Compare synchronous indexing, CDC pipelines and read-after-write patching at query time.
2 min answer -
advanced
A downstream team needs to react to order changes. The order service can publish events, or they can consume CDC from its database. Which, and why?
2 min answer -
advanced
A mobility platform wants operational database changes available to analytics and search within seconds, without adding load to the transactional databases. Design the change-data-capture pipeline and identify its principal failure modes.
3 min answer -
advanced
A team plans to publish database row changes via CDC as the organisation's domain events. What is wrong with that and what would you propose?
3 min answer
5 terms in this topic
CDC Initial Snapshot
The consistent full copy taken when a CDC pipeline starts, before streaming begins — and the step that determines whether the target is correct.
patternChange Data Capture in Practice
Turning a database's replication log into an event stream, so downstream systems learn about changes without the application publishing them.
case-studyLinkedIn Databus: Change Capture as a Product
LinkedIn built a change capture system so that derived stores — search, graph, caches — could stay current without every application dual-writing to them.
patternLog-Based CDC
Capturing changes by reading the database's own write-ahead log, which sees every change with no load on the source and no application involvement.
patternQuery-Based CDC
Detecting changes by repeatedly querying for rows modified since the last run — simple, universally available, and lossy in specific ways.
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.
Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.