Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
5 to work through
-
advanced
A data platform serves many internal teams with very different data volumes and sensitivity. How should tenancy be structured?
2 min answer -
advanced
A departing tenant's contract requires deletion within 30 days. Their rows sit in a shared 40 TB events table in an open table format - partitioned by event_date with tenant_id as one column of 60, snapshot retention 90 days, compaction weekly. An engineer runs a DELETE for that tenant_id. What happens next, and what is the 30-day clock actually measuring?
3 min answer -
advanced
A product-analytics platform serves many customers' data on shared infrastructure. What isolation is required at the storage and query layers?
2 min answer -
advanced
Pinterest described in 2015 how every object it stores carries a 64-bit identifier with the shard number packed into the high bits, so any object routes to its database without a lookup. Your analytics platform holds 40 TB of event tables for 900 internal tenants keyed by a random UUID, and an engineer proposes adopting the same trick for the warehouse. What does packing the tenant into the key actually buy on the analytical side, and where would copying Pinterest be a mistake?
3 min answer -
advanced
Your multi-tenant analytics platform stores all tenants in shared tables with a `tenant_id` column. What is the risk and how do you close it?
2 min answer
3 terms in this topic
Compute Isolation Boundary
The line across which one domain's analytical workload cannot affect another's performance, cost attribution or access.
conceptData Platform Tenancy
How a shared data platform separates domains and customers across storage, compute, catalogue and access control.
patternTenant-Scoped Layout Key
Making the tenant identifier the column that governs a table's physical organisation, so isolation, per-tenant cost attribution and per-tenant deleti…
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.