advanced 3 min answer

A lakehouse table takes 400 GB a day from a streaming job that commits every 60 seconds across 24 partitions. Daily compaction rewrites yesterday's data and a weekly job re-sorts the last 30 days. Roughly how much compute does that maintenance need, what share of the bill is it, and which assumption dominates the error?

compactionmaintenanceplanning-numberscostlakehouse
Show the full answer Hide the answer

The assumptions, stated

  • Files produced: 1,440 commits a day times 24 partitions is about 34,560 files a day, averaging about 12 MB. That is the thing compaction exists to remove.
  • Rewrite throughput: a compaction task reads, decodes, sorts and re-encodes. End to end this lands on the order of 30-60 MB of input per second per core for columnar files with a general-purpose codec. Take 40 MB/s as the central estimate and treat it as an estimate.
  • Compaction is read plus write, so one pass over a dataset costs twice its size in I/O.
  • Compute price: roughly $0.05-0.15 per core-hour at 2025 cloud rates, spot through on-demand.

The arithmetic

  • Daily compaction: 400 GB in, 400 GB out = 0.8 TB of rewrite I/O a day.
  • Weekly re-sort of 30 days: 12 TB in, 12 TB out = 24 TB a week, which is about 3.4 TB a day amortised.
  • Total: about 4.2 TB of rewrite I/O a day.
  • At 40 MB/s per core: 4.2e6 MB / 40 = about 105,000 core-seconds, which is roughly 29 core-hours a day, or about 880 core-hours a month.
  • Money: about $45-130 a month of compute.
  • Storage requests: 34,560 GETs and about 1,600 PUTs a day. At 2025 list prices for a major object store that is under $0.05 a day - noise.

The number and its range

Call it 15-60 core-hours a day and under $200 a month. On a platform whose query and pipeline compute is several hundred core-hours a day, maintenance is roughly 1-5% of the bill.

Which assumption dominates the error

The re-sort window, by a wide margin. Dropping the weekly re-sort from 30 days to 7 takes the total from 4.2 TB to about 1.4 TB of daily I/O - a 3x swing. Per-core throughput, the number people argue about, moves the answer by well under 2x. Sizing effort should go into deciding how far back the sort is worth maintaining, not into benchmarking the codec.

What the number rules in or out

  • It rules out the cost objection. "We cannot afford compaction" is almost never true. The real reasons tables go uncompacted are that nobody owns the job and it loses its capacity to deadline-bearing pipelines.
  • It rules in reserved capacity. A few tens of core-hours a day is cheap enough to give maintenance its own guaranteed window rather than leaving it as low-priority work that never runs.
  • It rules out re-sorting history nobody filters on. If 95% of queries cover 7 days, the other 23 days of re-sort buy nothing and are three quarters of the maintenance bill.

When this is the wrong answer

When the cost is not compute but the commit. A table with concurrent writers can have compaction fight the streaming job for the commit, because both must atomically swap the same file set, and a losing compaction retries from scratch. Past a few writers the constraint is commit conflict rate rather than CPU, and the arithmetic above says nothing useful about it. Estimate conflicts instead, and choose partition-scoped compaction windows - compact only partitions the streaming job has finished writing - rather than adding cores, which makes conflicts more frequent, not fewer.