File Formats & Compaction
Columnar formats, the small-file problem, and the maintenance nobody schedules.
6 to work through
-
beginner Multiple choice
A team keeps three years of event data as gzipped CSV in object storage because it is simple and anything can read it. The table is now about 500 GB compressed and the daily dashboard query reads four of the 60 columns. What has the simplicity actually cost them?
2 min answer -
intermediate
A lakehouse team raises its compaction target file size from 128 MB to 1 GB and switches the write codec from Snappy to Zstandard. Query cost falls. What did they give up, and when does that bill arrive?
2 min answer -
intermediate
A platform's analytical queries slow steadily over months with no change in data volume per day. What is happening?
2 min answer -
intermediate
A table written by a streaming job has become unusably slow. It holds 400 GB across 8 million files. Diagnose and fix.
2 min answer -
advanced
A lakehouse table takes 400 GB a day from a streaming job that commits every 60 seconds across 24 partitions. Daily compaction rewrites yesterday's data and a weekly job re-sorts the last 30 days. Roughly how much compute does that maintenance need, what share of the bill is it, and which assumption dominates the error?
3 min answer -
advanced
Review this configuration. A 60 TB events table is partitioned by ingest_date and written by a streaming job that commits every 60 seconds. A compaction job runs hourly over the last 24 hours with a 512 MB target and sorts each file by event_timestamp. Snapshot expiry runs monthly with 90-day retention. 85% of queries filter on customer_id over a 7-day range. What would you remove, what would you change and what would you leave alone?
3 min answer
3 terms in this topic
Compaction
Rewriting many small data files into fewer larger ones, with sorting, as a required background process - because streaming ingestion produces small f…
practiceRow Group Sizing
Choosing how many rows a columnar file groups into one statistics-bearing unit, which sets both the finest granularity a query can skip and the memor…
conceptSmall File Problem
The severe query degradation caused by a table stored as very many tiny files, where per-file overhead dominates actual data reading.
Neighbouring topics
Data Platform Architecture
General material on designing the analytical data estate end to end.
Medallion Architecture
Bronze, silver and gold layers, and what each layer is allowed to guarantee.
Open Table Formats
Iceberg, Delta and Hudi — transactions, snapshots and time travel over object storage.
Warehouse, Lake & Lakehouse
Three answers to where analytical data lives, and the workloads that separate them.
Storage Layout & Partitioning
Partition keys, clustering, and the scan the query planner is left able to skip.
Ingestion Patterns
Full load, incremental, append-only and merge, and the source system each one suits.
CDC Pipeline Design
Building on a change stream: snapshot plus delta, tombstones, and merge into the target.
Batch Orchestration
DAGs, dependencies, retries, and the difference between a schedule and an orchestration.
Workflow Schedulers
Airflow, Dagster and their kin — where the control plane sits and what it can recover.
Transformation Frameworks
Declarative SQL transformation with tests, lineage and versioned models.
Dimensional Modelling
Facts, dimensions, grain, and the star schema's continued relevance.
Data Vault Modelling
Hubs, links and satellites, and the auditability and load parallelism they buy.
Slowly Changing Dimensions
Overwriting, versioning or timestamping attribute history, and the reporting each enables.
Analytics Cost Control
Scanned bytes, idle warehouses, and the query nobody knew was running hourly.
Workload Isolation
Keeping an analyst's query off the pipeline's compute, and both off the dashboard's.
Data Platform Tenancy
Multiple domains on shared storage and compute, with separable access and cost.
Reverse ETL
Pushing modelled analytical data back into operational systems, and who owns it then.
Data Virtualisation
Querying across sources without moving data, and the performance ceiling that imposes.
Warehouse Migration
Moving off a legacy warehouse with thousands of reports pointed at it.