concept

Semantic Schema Break

also called Wire-Compatible Meaning Change

A change that keeps a field's wire format valid while changing what its values mean, so every compatibility check passes and every consumer computes a wrong answer without raising an error.

schema-registrycompatibilitydata-qualitycontractssilent-failure

A producer team changes amount from cents to whole currency units. The field is still an int. The schema registry's compatibility check compares types, names, defaults and nullability, finds nothing changed, and approves it. Every consumer continues to deserialise successfully and every downstream total is now 100 times too small.

This is the class of change no automated gate in the standard streaming toolchain can see. Compatibility checking is a structural comparison of two schemas; meaning lives outside the schema, in the shared understanding of the teams, and nothing in the pipeline holds a copy of it.

The symptoms are worse than a deserialisation failure, which stops a consumer in seconds with a clear exception. A semantic break produces plausible numbers that flow into dashboards, models, invoices and settlements until a human notices a figure that looks wrong.

Why it matters

Event streams are consumed by teams the producer has never met, and the retained history makes the mistake permanent. Records written before and after the change are indistinguishable from each other, so there is no predicate that finds the boundary unless someone recorded when the deployment happened.

The cost scales with the number of consumers and with retention. A 2-year ledger topic read by 40 teams means 40 independent wrong answers and a repair that has to be coordinated across all of them.

Implementation patterns

  • Treat a change of meaning as a new field name. amount_cents stays, amount_minor_units is added, both are populated during a dual-write window, and the old field is retired on a published schedule. This is the whole discipline in one rule, and it costs one deployment per producer.
  • Put units, currency, timezone and enum version in the field name or in a sibling field. price is a bug waiting to happen; price_usd_cents is self-describing to a consumer that only reads the schema.
  • Never reuse an enum value's numeric slot. Adding CANCELLED_BY_MERCHANT as a new value is safe; redefining value 3 from CANCELLED to REFUNDED is invisible to every check and changes the meaning of years of history.
  • Never reuse an id namespace. Recycling customer ids after deletion makes two different subjects share a key, which breaks every keyed aggregate and every compacted table.
  • Field-level statistical monitors on the consumer side: alert when a numeric field's p50 moves by more than a set factor, or when an enum's value distribution shifts sharply, inside one deployment window. This is the only automated detection that works, and it is detection rather than prevention.
  • A golden-record contract test that replays stored historical records through the new consumer and asserts expected outputs, so a change in interpretation fails a build.

Industry example

The gap is structural in the tooling rather than a failing of any one organisation. Confluent's Schema Registry, in production across the industry since 2015, implements BACKWARD, FORWARD and FULL compatibility as schema-to-schema structural comparisons; Protobuf and Avro both define compatibility in terms of wire encoding. None of them has a representation for "this integer is cents".

The recurring production shape is a marketplace or payments platform where a currency, unit or rounding convention changes during an expansion into a new country, the registry approves it, and the discrepancy surfaces in a finance reconciliation weeks later with a per-country pattern that nobody can explain from the logs.

Failure scenarios

  • Unit change. Cents to units, milliseconds to seconds, metres to kilometres. Totals are wrong by a constant factor, which looks like a data volume change rather than a bug.
  • Enum reinterpretation. A status value's meaning shifts, so a state machine downstream takes the wrong branch for a subset of records.
  • Timezone or clock source change. Event timestamps move by hours, which shifts every window boundary and produces an apparent traffic pattern change.
  • Nullability semantics. A field that previously meant "no value recorded" when absent starts meaning "explicitly zero", and averages change because the denominator changed.
  • Id namespace reuse. Two subjects share a key, so compacted tables hold a blend of both.
  • Meaning moved into an unstructured map. When governance makes schema changes slow, teams write new data into a generic string map, which has no schema at all and so cannot even be checked structurally.

Trade-offs

Choose Gains Pays
New field name per meaning change No silent breaks, history stays interpretable Field count grows, dual-write window, a retirement process
Reinterpret the existing field One deployment, no consumer work Every consumer silently wrong, history permanently ambiguous
Statistical monitors on consumers Detection within a deployment window Alert tuning, false positives on genuine business shifts

When not to use it

A topic with exactly one consumer owned by the same team does not need this ceremony. Change the meaning, change the consumer, deploy them together, and move on. The discipline exists because of the consumers you cannot coordinate with, so its weight should scale with the number of independent readers and the retention window, not be applied uniformly to every topic in the cluster.

Equally, a topic retained for 7 days and read only for live dashboards can absorb a reinterpretation with a note in a channel. The rule hardens where history is read: anything replayed to rebuild a store, or read by finance or a regulator, gets the new-field-name rule without exception.

Interview question

Q: Your schema registry is in FULL compatibility mode and a consumer team reports that their revenue figures dropped by two orders of magnitude overnight. The registry shows a schema change was accepted yesterday. Walk me through what happened and what you change.

What a strong answer covers: the registry compares structure, so a type-preserving change of units passes every check; confirm by diffing the two schema versions and looking for a field whose type is unchanged, then checking the producer's commit rather than the schema. The repair has two halves: a transform applied at read time for records in the affected offset range, which requires knowing the deployment timestamp, and a decision about whether to rewrite history or to carry a conversion in every consumer forever. The structural fix is the new-field-name rule plus a field-level distribution monitor that would have caught this in minutes rather than a day. A strong answer also notes that FULL mode gave the team false confidence, which is part of why the change went out unexamined.

Quick check

Quiz: Which class of schema change passes every compatibility mode and breaks every consumer? One that preserves the wire format while changing what the values mean, such as a unit change or an enum reinterpretation.

Flashcard: What is the single rule that prevents semantic schema breaks? A change of meaning is a new field name, never a reinterpretation of an existing one, with both populated during a published dual-write window.