Data Lakes & Lakehouses
Open formats on object storage with transactional metadata on top.
4 to work through
-
intermediate Multiple choice
A company is building a new analytics platform. Moderate data volume, heavy BI usage, strong governance requirements, a small data team. Warehouse or lakehouse?
2 min answer -
intermediate Multiple choice
A data platform runs interactive analyst queries alongside multi-hour batch pipelines on shared compute. Interactive latency becomes unpredictable. What is the architectural fix?
2 min answer -
advanced Multiple choice
A recommendation platform stores petabytes of interaction events used for both ad-hoc analysis and model training, with occasional need to delete individual users' data. Which storage architecture fits, and why does the deletion requirement drive the choice?
2 min answer -
advanced
You are asked to choose between a cloud data warehouse and a lakehouse for a new analytics platform. How do you decide?
2 min answer
4 terms in this topic
Lakehouse
Open table formats over object storage that add transactions, schema enforcement and incremental updates to a data lake.
conceptSchema Evolution
Changing a table's structure over time while keeping existing data readable and existing consumers working.
conceptTime Travel
Querying a table as it existed at a previous version or timestamp, made possible by keeping the metadata and files of prior commits.
patternWorkload Class Isolation
Running interactive, scheduled and batch analytical workloads on separate compute over shared storage, so that a long-running job cannot make an anal…
Neighbouring topics
Data Architecture
General material on structuring, storing and governing data.
Relational Modelling
Normalisation, keys, constraints and the invariants a schema enforces.
NoSQL Stores
Key-value, document, wide-column and graph — what each buys and forbids.
Indexing
Designing indexes per query shape, and paying for them on every write.
Query Optimisation
Reading a plan, fixing statistics, and finding the real bottleneck.
Transactions & Isolation
ACID, isolation levels, and the anomalies each level permits.
Replication
Primaries, replicas, lag, and synchronous versus asynchronous durability.
Partitioning & Sharding
Splitting data across machines, and the one-way door of a partition key.
Caching Strategies
Cache-aside, read-through, write-through and where each belongs.
Cache Invalidation
Stampedes, penetration, staleness windows and versioned keys.
CQRS
Separating the write model from the read models that serve queries.
Event Sourcing
Storing the change log as the system of record, and what that costs forever.
Change Data Capture
Turning a database's replication log into a stream, and its coupling risk.
Data Warehousing
Dimensional modelling, star schemas and analytical workloads.
ETL & ELT
Where transformation happens, and how much raw history you keep.
Streaming Data
Windowing, watermarks, late arrivals and exactly-once semantics.
Data Governance
Ownership, lineage, quality, catalogues and who may see what.
Data Lifecycle & Retention
How long data is kept, where it ages to, and how it is actually deleted.
Polyglot Persistence
Choosing a store per workload, and the operational cost of variety.