Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
5 to work through
-
advanced
A bank implemented Data Vault for auditability. Analysts say it is unusable and are extracting to spreadsheets. What went wrong?
2 min answer -
advanced
A team proposes Data Vault for a customer domain. The source is one table of 40 columns, of which 6 are business or foreign keys and 34 are descriptive attributes changing at three clearly different rates. Before agreeing, estimate how many physical tables the vault produces, how many joins a plain "current customer with address and account status" query needs, and how many rows five years of history holds for 8 million customers. Which assumption dominates the error?
3 min answer -
advanced
An interviewer says - a regulator can ask us what any published report said on any date in the past two years and why it said that. Our platform already has a Data Vault. Where do you take this?
3 min answer -
advanced
Review this estate. Postgres and three SaaS sources land raw in bronze. A Data Vault of hubs, links and satellites is built in silver. Star schemas are built in gold. Six departmental extracts are then copied into departmental warehouses. Freshness is 40 minutes, the team is six people, and there are 42 dashboards. What would you remove, what would you change, and what would you leave alone?
2 min answer -
advanced
What problem does data vault modelling solve, and when is its complexity justified?
2 min answer
3 terms in this topic
Business Key Collision
Two different real-world entities merging into one hub row because the key chosen for it is unique only inside one source system - a corruption that …
patternData Vault Modelling
A modelling approach separating business keys, relationships and descriptive attributes into hubs, links and satellites, optimised for auditability a…
patternHub and Satellite
Separating stable business keys from their changing attributes and from their relationships, so each can be loaded independently and kept forever.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.