Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
5 to work through
-
intermediate
A data platform's interactive queries become unpredictable whenever a large pipeline runs. What is the fix, and what remains shared?
2 min answer -
intermediate Multiple choice
A lakehouse platform runs interactive BI on its own compute cluster and pipelines on another, and both are sized comfortably. Every Tuesday and Wednesday queries against the ten largest tables run 3 to 6 times slower on both clusters at once, then recover by Thursday. Table maintenance - compaction and snapshot expiry - is submitted as low-priority work on the pipeline cluster. Which change fixes it?
3 min answer -
intermediate
Every morning at 09:00 executive dashboards are slow. Investigation shows a data scientist's exploratory query. How do you fix this permanently?
1 min answer -
advanced
A data platform's ad-hoc queries, scheduled pipelines and machine-learning training compete for the same capacity. How should they be isolated?
2 min answer -
advanced
An analytics platform separates storage from compute so many teams can query the same data. How should result caching, per-team compute isolation, concurrency scaling and cost attribution be designed — and what workloads defeat the cache?
3 min answer
3 terms in this topic
Table Maintenance Budget
Treating compaction, clustering and snapshot expiry as a workload class with its own reserved compute and a guaranteed window, rather than as low-pri…
conceptWarehouse Concurrency Scaling
Adding compute clusters to absorb concurrent queries rather than queueing them, and the cost behaviour that turns a queue into a bill.
practiceWorkload Isolation
Preventing one team's expensive query or backfill from degrading everyone else's analytics, by separating compute rather than sharing one pool.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.