concept

Data Product

also called Published Dataset, Owned Data Asset, Analytical Interface

A dataset published deliberately as an interface with a named owner, documented semantics, quality guarantees and a change policy - as distinct from a table that consumers happen to be able to reach.

data-meshcontractsownershipdiscoverabilityquality

Most analytical consumption reads whatever tables exist: a service's raw ingested database copy, an intermediate table from someone else's pipeline, a view built for a different purpose. None of these were published; they were merely reachable, and their owners made no commitment about them.

A data product inverts that. It is a dataset produced deliberately for consumption, carrying:

  • A named owner, accountable for it.
  • A defined schema, enforced at publication.
  • Documented semantics — what each field means, its grain, its units, its time basis, what null indicates.
  • Quality guarantees stated as measurable expectations: freshness, completeness, key uniqueness, value ranges.
  • A change policy: additive changes freely, breaking changes with notice and a migration path.
  • Discoverability, so consumers can find and understand it without asking a person.

Why it matters

The distinction between "published" and "reachable" is the whole thing. A consumer reading a reachable table has created a dependency the producer cannot see, on an object the producer will change without warning — and when it breaks, the producer's change was correct, the breakage is in someone else's system, and neither party had the information to prevent it.

Semantics matter more than schema. A schema change breaks loudly and is fixed in a day. A semantic change — a field's meaning shifting, a status value repurposed, a currency changing — produces wrong answers silently, and is discovered when someone questions a number weeks later.

Implementation patterns

  • Published separately from internal storage. The product is a deliberate artefact, not the producer's operational table — the same reasoning as an outbox event rather than raw change capture, and it is what lets the producer refactor freely.
  • Schema enforced at publication, so violations fail at the source rather than downstream.
  • Quality expectations monitored and alerted on, so a violation is detected by the producer rather than reported by the consumer.
  • Additive-only evolution by default, which is what keeps the arrangement affordable.
  • A consumer registry, which is what makes a notice period actionable — and which the producer needs in order to know a change is safe.
  • Service-level expectations stated: how fresh, how available, how supported.
  • Applied selectively. Contracts belong on the datasets that cross team boundaries and matter, not on every table — universal application is unaffordable and dilutes the signal.
  • A quality gate that blocks propagation, holding publication and serving the previous good version rather than passing bad data downstream with an alert attached.

Industry example

Data products are the central unit of the data mesh model, where domain teams publish their analytical data as products supported by a central self-service platform. The same idea appears independently as data contracts and as schema-registry compatibility enforcement in streaming platforms — all three are the recognition that anything a consumer can reach becomes a contract whether or not anyone intended it.

The organisational precondition is consistently the deciding factor: a domain team that publishes a product and does not maintain it has made things worse than the raw table it replaced, so accountability must be real rather than declared.

Failure scenarios

  • A product that is the internal table renamed, which formalises the coupling without removing it.
  • Schema documented, semantics not, so meaning changes silently.
  • Quality expectations written and unmonitored.
  • No consumer registry, making notice periods meaningless.
  • Products declared without ownership capacity, which is worse than no product.
  • Contracts on everything, which is unaffordable and never completes.
  • The producer bearing all the cost and receiving none of the benefit — which is why "you may change anything not in the contract" is essential to getting agreement.
  • No quality gate, so bad data propagates with an alert rather than being stopped.

Trade-offs

Publishing a product imposes real ongoing cost on the producer: a separate artefact to maintain, quality to monitor, notice periods to honour, and consumers to consider before making changes they would otherwise make freely.

It slows the consuming side too, at least initially — requesting a field rather than joining to a table is worse for the consumer when the producer is unresponsive, and that friction is a genuine adoption obstacle.

The trade is producer effort and consumer convenience in exchange for a dependency both parties can see and manage. For a small organisation with a handful of pipelines, direct reads and good communication are proportionate. At the scale where nobody can enumerate the consumers of a table, the implicit arrangement has already failed — the breakages are simply being absorbed as background noise and attributed to the data platform.

Interview question

"We want to make our order data a data product. Tell me what specifically we are committing to, what we get in return that makes it worth doing, and what you would refuse to put in the contract."