A departing tenant's contract requires deletion within 30 days. Their rows sit in a shared 40 TB events table in an open table format - partitioned by event_date with tenant_id as one column of 60, snapshot retention 90 days, compaction weekly. An engineer runs a DELETE for that tenant_id. What happens next, and what is the 30-day clock actually measuring?
Show the full answer Hide the answer
Second by second
The planner looks for files containing matching rows. Because the tenant's events arrived interleaved in time with everyone else's, nearly every file in every partition they were active in holds at least one of their rows. With copy-on-write deletes, each of those files is rewritten without the matching rows.
That is a rewrite on the order of the whole table: 40 TB read and roughly 40 TB written. At a few tens of MB/s per core it is hundreds of core-hours - a few hours on a large cluster - and for the duration storage roughly doubles, because the old files cannot be removed yet.
Then the commit lands, a new snapshot becomes current, and the tenant's rows are still there. The previous snapshots still reference the old files, which is what time travel means. Anyone with table access can read the deleted tenant's data for the next 90 days.
Where it amplifies
- Downstream incremental models that pick up changed files see most of the table as changed. Either they reprocess three years of history tonight, or they are keyed on append-only semantics and skip the deletion entirely, leaving the tenant's rows alive in every derived table.
- Merge-on-read looks like the fix and is not. A delete vector commits in seconds, but the bytes stay until compaction rewrites the file, and every read of that file now pays a merge. You have traded a visible 4-hour job for an invisible tax and the same residual data.
- Copies nobody listed: the nightly backup, the cross-region replica, the BI tool's extract, the CSV an analyst saved in 2024.
What the clock is actually measuring
Not the DELETE. The clock measures commit time plus snapshot expiry interval plus compaction interval plus backup retention. With 90-day snapshots and weekly compaction, the honest answer to the tenant is about 91 days, and the contract says 30.
What stops it
- Make the tenant a layout property, not a column. Partition or sort by tenant so their data is a known file set and deletion is a metadata operation. Datadog has described doing this in Husky for isolation reasons - data from two tenants is never mixed into one file (2023 engineering blog) - and the same layout is what makes per-tenant deletion and per-tenant cost attribution tractable.
- Shorten retention where there is a deletion duty. 7-day snapshots on tables carrying customer data, 90 elsewhere.
- Make expiry part of the deletion job, not a separate monthly chore, and keep its completion log as the evidence of deletion.
- Maintain the copy inventory. The deletion runbook names every replica and backup, or the certificate you sign is false.
When this is the wrong answer
When the table is small or the tenant is bounded in time. A 200 GB table, or a tenant whose rows live in three monthly partitions, is a DELETE plus an immediate expiry - minutes of work. Re-laying out a 40 TB table for one deletion a year is the expensive mistake in the other direction; what you need then is a short retention setting and a documented procedure.