Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
5 to work through
-
intermediate
A company wants warehouse data pushed back into operational systems. What are the design considerations, and what should not go through this path?
2 min answer -
intermediate
A platform pushes warehouse-derived data back into operational systems. What are the risks, and what must be designed?
2 min answer -
intermediate
Marketing wants warehouse segments pushed into the campaign tool hourly. What do you require before agreeing?
2 min answer -
advanced
A warehouse table now feeds the CRM through reverse ETL. It breaks, and the analytics team and the CRM team each say it is the other's problem. How do you resolve it?
1 min answer -
advanced
At 01:50 an upstream system renames a column and the nightly customer model's join silently yields 4,000 rows instead of 1.2 million. Every test on that model is a not-null check on columns that survived, so the model passes. At 02:30 reverse ETL syncs it to the CRM. By 09:15 sales reports that lifecycle stage and account owner are blank on more than a million accounts and overnight campaigns fired against the wrong segment. Nothing errored. What failed, and which design decision made it possible?
3 min answer
3 terms in this topic
Absence-As-Delete Sync
A synchronisation mode that treats a record's absence from the source query as an instruction to remove or clear it downstream - so any query returni…
patternOperational Data Push
Sending modelled analytical data back into operational tools, which turns a warehouse table into a production dependency with none of the guarantees.
patternReverse ETL
Pushing modelled data from the warehouse back into operational systems so that business tools act on the same definitions analysts report on.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.