Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
5 to work through
-
intermediate
An analytical platform's cost has grown faster than its usage. Which controls actually reduce it?
2 min answer -
intermediate
An analytics platform's cost is growing faster than its usage. What causes it, and which controls work without blocking analysts?
2 min answer -
intermediate
One analyst's query cost £4,000 in a single afternoon. Leadership wants a policy. What do you actually implement?
1 min answer -
intermediate
The analytics bill has tripled in six months with no corresponding growth in users or data volume. Where do you look?
1 min answer -
advanced
An analytics platform's query costs are growing faster than its data. What are the drivers and controls?
2 min answer
4 terms in this topic
Analytics Cost Control
The set of design and operational choices that determine whether an elastic data platform costs a predictable amount or an alarming one.
conceptBytes Scanned
The volume of data a query reads - which in consumption-priced analytical systems is both the performance limit and the bill, making layout a cost control.
patternQuery Cost Ceiling
A pre-execution limit on the bytes or credits a single query may consume - converting an unbounded bill into a failed query the author can see and fix.
metricScanned Bytes
The volume a query reads, which is what most analytical engines bill for and what almost every optimisation ultimately reduces.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.