practice

Table Maintenance Budget

also called Maintenance Capacity Reservation, Housekeeping Window

Treating compaction, clustering and snapshot expiry as a workload class with its own reserved compute and a guaranteed window, rather than as low-priority work that loses to every deadline.

compactionsnapshot expiryisolationsmall filescapacity

Interactive analytics has its own cluster. Pipelines have another. Both are sized comfortably, and every Tuesday and Wednesday queries against the largest tables run several times slower on both at once, then recover. Nothing in the isolation design predicted a correlated slowdown, because the design isolated compute and the tables are shared.

Table maintenance is the third workload, and it is the one nobody funds. Compaction, sorting, clustering and snapshot expiry are submitted as low-priority background work on a cluster whose foreground work has a published deadline, so maintenance loses every time the deadline is close. A table maintenance budget is the decision to give that work its own capacity and its own window, with the same status as any other scheduled job.

Why it matters

The cost of unperformed maintenance is per-file, and per-file cost is invisible to the isolation model that separates compute. Opening a file on object storage costs on the order of 20 to 60 ms, and a query planner lists and reads statistics for every candidate file before a row is read. A table drifting from 5,000 files to 200,000 has multiplied its planning work by 40 while its data volume barely changed, and every reader pays regardless of which cluster they are on.

It degrades on a curve nobody watches. Job success rates stay at 100%, queries keep returning correct answers, and the slope is gentle enough that each week feels like the last. The usual discovery event is a bill or a complaint, months after the drift began.

Implementation patterns

  • Reserve capacity, do not borrow it. A small dedicated cluster or a reserved slot pool that maintenance always gets, so it is not competing with a deadline it will always lose to.
  • A guaranteed window per table class, sized from the write pattern: streaming tables need frequent compaction of the current partition, batch tables need one pass per completed period.
  • Compact once per period, not repeatedly. Hourly compaction across a rolling 24 hours rewrites yesterday's data 24 times, which is pure write amplification.
  • Alert on the physical signal, not the job. Files per partition, average file size and snapshot count are the metrics; "compaction succeeded" tells you nothing about whether it kept up.
  • Expire snapshots on a cadence matched to the rewrite rate. A table whose nightly merge rewrites most of its files and retains 90 days of snapshots stores many multiples of its logical size and gives the planner a large metadata tree to walk.
  • Charge maintenance to the table's owner in cost reporting, so a team that chooses minute-level commits sees the housekeeping that choice creates.

Industry example

The large observability vendors that publish their event-store designs converge on the same shape: writers, readers and compactors as independently scaled services over shared object storage. The separation is the point. The compactor is not a background thread inside the writer competing for the writer's resources, it is a service with its own capacity, scaled for its own work. Most teams operating a lakehouse in 2024 have the same three roles and only two of them funded, which is why the symptom shows up as a weekly rhythm rather than as a failure.

Failure scenarios

  • Correlated degradation across isolated clusters, which sends the investigation towards compute and away from layout.
  • Maintenance that can never catch up. Once a table is far enough behind, a compaction pass costs more than the window allows, so it fails or is killed and the deficit grows every day.
  • Monthly expiry that times out. Deferring expiry produces one enormous deletion job that exceeds its runtime, so retention is never actually enforced and storage grows without limit.
  • Compaction fighting the writer. Maintenance and a streaming writer commit to the same table concurrently, and one side's commit is rejected and retried, wasting the window on conflict.

Trade-offs

Choose Gains Pays
Dedicated maintenance capacity Predictable query performance; layout never drifts Capacity idle much of the week, roughly a small always-on cluster
Maintenance on the pipeline cluster at high priority No extra spend Starves the workload whose freshness commitment is published
Maintenance at low priority No extra spend and no visible conflict Loses every busy period, which is exactly when the files accumulate

When not to use it

An append-only table written by hourly batch jobs in files of a few hundred megabytes needs almost no maintenance, and reserving capacity for it is spending money on an idle cluster to solve a problem the write pattern already prevents. The same is true of small tables, where a rewrite is seconds of work that fits anywhere. Maintenance earns a budget when writers commit frequently, when merges rewrite existing files, or when the table is large enough that catching up is not possible inside a normal job slot - which in practice means streaming ingestion and update-heavy change data capture.

Interview question

Q: Your lakehouse has separate compute for BI and for pipelines, and both slow down together twice a week. Walk me through how you would confirm the cause, and what you would change before buying any more compute.

What a strong answer covers: recognising that simultaneous degradation of isolated clusters points at shared storage rather than shared compute; asking for files per partition and average file size over time rather than CPU; connecting the weekly rhythm to a maintenance job that loses priority during the peak; proposing reserved capacity and a guaranteed window instead of larger clusters; and naming the alert that makes the failure visible next time, which is a physical layout metric rather than job success.

Quick check

Quiz: Why does separating BI and pipeline compute fail to protect either one from a compaction backlog? - Because the separation isolates CPU while both clusters read the same files, and the cost of a degraded layout is per-file planning work that every reader pays.

Flashcard: Which metric tells you maintenance is falling behind, and which one will not? - Files per partition and average file size will; compaction job success rate will not, because the job succeeds at doing less than the table accumulates.