Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
6 to work through
-
beginner Multiple choice
A 40-person company has 80 GB of analytical data growing about 4 GB a month, six analysts, a nightly load from Postgres and four SaaS sources, and no data engineer. An architect proposes an object-storage lakehouse with an open table format so the company never has to migrate again. What should they build first?
2 min answer -
intermediate
How should a platform decide between a warehouse, a lake and a lakehouse for a given workload?
2 min answer -
intermediate
Notion runs its product on sharded Postgres and originally loaded analytics into a managed warehouse through off-the-shelf connectors. In 2022 it built its own data lake on change data capture into Kafka, Hudi tables and S3 instead. What forced that, and where would copying it be a mistake?
2 min answer -
advanced
A retail client needs BI dashboards for 400 concurrent users and a data science platform over the same data. One lakehouse or a lakehouse plus a serving layer?
2 min answer -
advanced
Datadog has published the design of Husky, its third-generation event store, which splits writers, readers and compactors into independently scaled services over a shared metadata store and commodity object storage rather than running one storage tier. What problem forces that separation, and where would copying it be a mistake?
2 min answer -
advanced
Leadership wants to consolidate a warehouse, a data lake and three departmental marts into a lakehouse. How do you scope and sequence this?
2 min answer
2 terms in this topic
Lakehouse
An architecture combining a lake's cheap open storage with a warehouse's transactions, schema enforcement and query performance, through a table form…
practiceWorkload Fit Assessment
Choosing between warehouse, lake and lakehouse by the workloads that must run, rather than by which one is currently fashionable.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.