Review this. A payments platform runs 300 topics with the schema registry in FULL compatibility mode for every subject, a rule that each new field must be optional with a default, and a schema-change request that goes to an architecture review board with a three-week turnaround. Engineers have started putting new data into a generic metadata map to avoid the queue. What would you remove, what would you change, and what would you leave alone?
Show the full answer Hide the answer
What is actually required
Two guarantees, with very different time horizons, and the design conflates them.
Old readers must not break on new data for as long as old readers exist, which is the rolling-deploy window: 2 to 48 hours in production. New readers must not break on old data for as long as that data is retained, which for a payments ledger is years and for a clickstream is 7 days. FULL compatibility satisfies both by forbidding nearly every change, and applies that to all 300 topics regardless of which horizon each one actually has.
What I would remove
The three-week review board, for the roughly 280 topics that are not ledger topics. It has already
failed, and the generic map field is the proof. An escape hatch engineers build for themselves is worse than
any change it avoided: a map of strings has no schema at all, so the registry cannot check it, CI cannot
test it, and a consumer reading metadata["amount"] has no idea what unit it is in. The governance process
produced exactly the outcome it was designed to prevent, at a cost of three weeks per change.
Replace it with two automated gates: the registry's compatibility check running in the producer's pipeline, and a consumer-side contract test that replays a stored set of real historical records through the new deserialiser. Both run in minutes and both catch more than a board meeting does.
The one change that matters
Set compatibility per subject rather than per cluster. Choose FULL for the ledger topics whose history is read years later, and BACKWARD for topics whose replay horizon is the 7-day retention, which permits adding a required field once consumers are upgraded. The compatibility mode is a statement about who upgrades first, and that answer is genuinely different for a two-year audit log and a seven-day clickstream. One cluster-wide setting cannot express it.
What I would leave alone
The "every new field is optional with a default" rule, which looks like duplication of the registry's job and is not. The registry checks that two schemas are compatible. It does not know that during a rolling deploy, an older producer is still writing records without the new field, and a reader whose schema marks it required will fail on those. The default value is what makes the deploy window survivable. Keep the rule and keep it boring.
I would add one rule the current design has no equivalent of: changing the unit, the enum meaning or the id namespace of an existing field is a new field name, never a reinterpretation of the old one. No compatibility check sees that class of change, because the wire format is identical and only the meaning moved. It is the most damaging schema change available and the registry will approve it instantly.
How I would argue this in the review
The board's cost is measurable and its benefit is already zero. Three weeks per change multiplied by the change rate is the price, and the map field means nothing is actually being reviewed. Offer the trade explicitly: the board keeps authority over the ledger subjects, where the two-year horizon makes a mistake expensive and irreversible, and gives it up everywhere else. Narrowing scope is easier to agree than removing a control.
When this is the wrong critique
If a regulator requires documented human review of changes to reported data structures, the board stays, and the engineering answer is to make its scope precise rather than to argue it away. The failure here is not that review exists; it is that review was applied uniformly to 300 topics with different risk, which guaranteed the queue and therefore the workaround.