Systems & Scale 4 October 2026 7 min read 1,642 words

A delete with no address

Apache Iceberg has prohibited new equality deletes in version 4 — the one write path that never had to know which file it was changing. The prohibition is forward-only, and every v4 reader must keep applying the deletes that no v4 writer is allowed to create.

The argument

Iceberg has prohibited the one write path that never read what it changed, and because the prohibition requires no rewrite, every reader must keep applying the deletes that no writer is allowed to create.

A deletion vector in an Apache Iceberg table knows exactly what it is deleting. So does a position delete file. Both carry a referenced_data_file — a path, and within it an offset. A reader planning a scan can look at one and decide, before opening anything, whether it has any bearing on the file in front of it.

An equality delete carries no address at all. It carries a predicate — id = 5 — and a sequence number, and the rule for applying it is that it hits every data file in the same partition with a lower sequence number. If it was written under an unpartitioned spec, the specification is blunt about the consequence: such deletes are "applied as global deletes," across the table. The row being removed might live in one file or in four hundred. Nothing in the delete says which, and there is no metadata to ask.

On 29 September a thirteen-line change landed in format/spec.md and stopped anyone writing new ones. Writers must not add equality delete files to v4 tables; equality deletes cannot be added as an entry to a v4 manifest. The summary of what version 4 brings is now two items long: relative paths in metadata, and "Writing new equality deletes is no longer allowed." A format version that is otherwise about where files are named has quietly withdrawn a capability.

The capability it withdrew is the only write path in Iceberg that never had to read.

That is the whole reason it existed. A change-data-capture stream arriving at tens of thousands of rows a second has to express "this row is now different." Done with a deletion vector, that requires first finding the row — which file holds it, at which offset — which means a lookup against object storage on the ingest path, before every write. Done with an equality delete, it requires knowing only the key. Iceberg's own Flink documentation still frames upsert as a version-2 feature: upsert is supported "based on the primary key when writing data into v2 table format," and the table must declare equality fields. The predicate delete is what allowed a streaming sink to stay ignorant of the table's physical layout. That is not laziness. It is the difference between a pipeline that keeps up and one that falls behind.

The usual complaint about this bargain is read amplification, and it is the least interesting thing wrong with it. Readers pay to join against delete files; compaction pays it back. Iceberg ships maintenance actions for exactly this, including a metadata-only pass that drops equality delete files whose sequence numbers put them behind every live data file in a partition. A deferred cost that can be settled on a schedule, with compute you were going to buy anyway, is a respectable trade.

The cost that cannot be settled is a different one. It is what the table becomes unable to say about itself.

Version 3 added row lineage: a _row_id assigned to every row and a _last_updated_sequence_number carried forward when it changes. The specification immediately carves out an exception, and states the reason plainly: lineage is not tracked for rows updated via equality deletes, "because engines using equality deletes avoid reading existing data before writing changes and can't provide the original row ID for the new rows." Such an update is recorded as the complete removal of one row and the arrival of an unrelated new one. The history of a customer record maintained this way is not a history. It is a sequence of strangers who happen to share a key.

The same gap shows up in arithmetic. Partition statistics include total_record_count, the number of rows after deletes are applied. For a table carrying equality deletes, computing it "requires reading data," so implementations "may omit this field and must write NULL, indicating that the exact record count in a partition is unknown." A table written this way cannot count its own rows from its own metadata.

There is a smaller irony in which mechanism survived. The specification notes that both position and equality delete files may encode the values of the deleted row, and that this "can be used to reconstruct a stream of changes to a table" — the older, cruder way of getting change data out of a lakehouse, which equality deletes support perfectly well. It is the newer and more precise instrument, row lineage, that they defeat. An architecture that wants to know what changed can still scrape it out of the delete files, the way it always did. An architecture that wants the table to tell it which row this is, and when it was last touched, cannot. The capability that was lost is not the ability to observe change; it is the ability to be told about it by the system of record rather than inferring it from the exhaust.

Read in sequence, the versions describe a rule the project never stated outright. Version 3 prohibited new position delete files in favour of deletion vectors; both of those name a file. The specification has not quite caught up with its own firmness — one section still calls position delete files deprecated while the list above it calls them prohibited — but the direction is not in doubt. Version 4 removes the one encoding that names nothing. The direction is not toward faster deletes. It is toward deletes that the table can account for — and a write that skipped the read is a write the table can only describe as an absence followed by a coincidence.

The strongest case against the prohibition is that it moves a real cost onto the systems least able to absorb it, and offers nothing in exchange. The word "upsert" does not appear in the table specification. Formats describe encodings and engines choose strategies, which is a defensible division of labour, but it means the capability has been removed in a document that takes no position on what should replace it, by people who are not the ones running the Flink job at three in the morning. The honest defence of equality deletes was never that they were free. It was that their cost was deferrable, bounded, and purchasable later with compute — which is a better position than unbounded write latency paid on the critical path, every row, forever. That case deserves more than it has been given. It should also be said that what is argued here is read off Iceberg's own specification, pull request and documentation; that record is authoritative about what the format now permits, and silent about what the change will cost anyone in production.

Then there is the part that makes the prohibition stranger than it first reads. It is not a removal. Upgrading a v2 or v3 table to v4 "does not require rewriting data or delete files," and readers "must continue to apply equality deletes for v2 and v3 tables and for equality deletes carried over into upgraded v4 tables." So the equality deletes already written stay where they are, inside tables that now declare themselves version 4, and every reader implementation in every language must carry the code to apply them indefinitely. The feature is dead on the write side and immortal on the read side. The clause that makes the upgrade painless is precisely the clause that makes the liability permanent. Nothing in the format will ever force the rewrite that would clear it; someone has to decide to pay for it, and the upgrade path is explicitly designed so that nobody has to.

It is worth noticing what the mechanism rested on all along. An equality delete matches rows by the table's identifier fields, and Iceberg is candid that "uniqueness of rows by this identifier is not guaranteed or required," being "the responsibility of processing engines or data providers to enforce." The format has been applying a key-based delete against keys it never checked were keys. A request that writers at least validate the column constraints the specification does impose — opened in May 2025 — was closed as not planned, with the reporter observing that violations surfaced only when the table was read. That is the shape of the whole arrangement: a write that checks nothing, and a read that discovers everything.

This is what makes a file format a different kind of artefact from an API. An API deprecation is a conversation with the people currently calling it; they migrate, or they stay behind, and in a few years the question is closed. A table format deprecation is a conversation with data that will outlive the engine, the team and quite possibly the company. The bytes in object storage are the long-lived thing, and a specification is the only document that still governs them once everyone who chose the ingestion architecture has moved on. Thirteen lines merged on a Tuesday will be shaping what streaming pipelines are allowed to do a decade from now, and the people affected by it mostly were not watching.

The deletes already written are not going anywhere. They are scattered across partitions, addressed to nothing, waiting to be applied by readers who will never be allowed to write one. Whatever you make of the prohibition, the part worth keeping is the diagnosis underneath it: a write that skips a read is not cheap, it is financed, and Iceberg has now published the terms. The next time a design offers to make writes free by not looking at what they change, the question to ask is not what it will cost to read afterwards. It is what the system will be permanently unable to say about itself — and who, in ten years, will still be paying the interest on that.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Spec: Forbid writing new equality deletes in v4 (PR #17783) Apache Iceberg · 2026-09-29
  2. Iceberg Table Spec, format/spec.md at main Apache Iceberg · 2026-10-04
  3. Commit 7ba28c3: Spec: Forbid writing new equality deletes in v4 Apache Iceberg · 2026-09-29
  4. Flink Writes documentation, docs/docs/flink-writes.md Apache Iceberg · 2026-10-04
  5. Maintenance documentation, docs/docs/maintenance.md Apache Iceberg · 2026-10-04
  6. Equality delete column constraints are not enforced (Issue #12971) Apache Iceberg · 2025-05-05

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

iceberglakehousestreaming ingestionwrite pathdeprecation