intermediate 3 min answer Multiple choice

A lakehouse platform runs interactive BI on its own compute cluster and pipelines on another, and both are sized comfortably. Every Tuesday and Wednesday queries against the ten largest tables run 3 to 6 times slower on both clusters at once, then recover by Thursday. Table maintenance - compaction and snapshot expiry - is submitted as low-priority work on the pipeline cluster. Which change fixes it?

compactionmaintenanceisolationsmall filesmetadata
Pick one
Show the full answer Hide the answer

The deciding property

Separating compute isolates CPU; it does not isolate the table. Both clusters read the same files, so anything that degrades the physical layout degrades both at once. That simultaneity is the diagnostic clue: two isolated clusters slowing together points at shared storage, not at shared compute.

The weekly rhythm is the second clue. Maintenance runs at low priority behind the Tuesday and Wednesday pipeline peak, loses, and small files accumulate for two days until Thursday's quieter window lets it catch up.

Why isolating maintenance is the answer

A query's cost has a per-file component that no amount of CPU removes. Opening a file on object storage costs on the order of 20 to 60 ms of latency, and the planner must list and read statistics for every candidate file before a single row is read. A table that drifts from 5,000 files to 200,000 has multiplied its planning work by 40 while its data volume barely moved, and that is why both clusters degrade together with idle cores.

Maintenance is therefore a third workload class with its own service commitment, not a background courtesy. Give it its own compute, a window it is guaranteed to get, and an alert on file count per partition rather than on job success.

Why the other options fail

  • Bigger clusters. The bottleneck is per-file latency and planning, not compute. Doubling both clusters buys a small improvement at double the cost, and the file count keeps growing, so the purchase has to be repeated.
  • Raise maintenance priority on the pipeline cluster. This works, by taking the capacity from the pipelines, whose freshness commitment is published to consumers. It moves the starvation to the workload with the promise attached, and it will be quietly reverted after the first missed feed.
  • Turn off snapshot expiry. The exact opposite move. Retained snapshots hold the pre-rewrite files alive, so storage grows and the metadata the planner walks grows with it. Planning gets slower and the storage bill rises.
  • Faster storage and a larger cache. Caching data files does not remove listing and planning work, the cache is cold at the start of each day, and the 200,000 small files defeat the cache's own block granularity. This is the option that sounds like it addresses the storage clue and addresses the wrong part of it.

What it costs

Dedicated maintenance capacity is idle for much of the week, so it is a real line on the bill, roughly the price of a small always-available cluster. The guaranteed window also means pipelines occasionally wait behind maintenance. That is the trade: a predictable small cost against an unpredictable large one that nobody attributes correctly.

When this is the wrong answer

If the tables are append-only and written by hourly batch jobs producing files of 200 MB or more, compaction has almost nothing to do and snapshot expiry is minutes of work. Dedicating capacity to it is waste, and the weekly slowdown has a different cause worth finding. Maintenance earns its own capacity when writers commit frequently or when merges rewrite files, which in practice means streaming ingestion or update-heavy change data capture.