advanced 2 min answer

Salesforce's Force.com platform was described by Weissman and Bobrowski at SIGMOD 2009 as metadata-driven: many tenants' records share physical tables, with metadata rather than per-tenant schemas describing what each column holds. Your analytics estate sits on a platform of that shape. What breaks in a column-level data catalogue, and what replaces it?

salesforcemulti-tenancymetadatacatalogueclassificationlineage
Show the full answer Hide the answer

The situation they were in

Per-tenant schemas mean per-tenant DDL, and DDL per tenant does not survive tens of thousands of tenants: every schema change becomes a fleet operation and the database's own catalogue becomes the bottleneck. The published design avoids that by keeping records in shared tables with generic columns, and holding the meaning of each column in a metadata dictionary. Tenants add fields without any DDL running.

What that does to a catalogue

Every column-level catalogue on the market assumes the physical column is the unit of meaning. Classification is attached to a column, lineage edges are drawn between columns, and glossary terms are linked to columns. On a metadata-driven store, one physical column holds a different tenant-defined field for every tenant, so all three of those bindings are attached to an object that does not have a single meaning.

The concrete failure is silent rather than loud. A harvester runs, succeeds, and reports a column called val3 with mixed content. Somebody classifies it Internal because most of the sampled rows look innocuous, and one tenant's health identifiers now sit under an Internal policy. Nothing errors.

What replaces it

The catalogue's unit becomes the pair of tenant and logical field, resolved through the metadata dictionary, not the physical column. That has three consequences worth stating in a design review:

  • Classification is carried on the logical field and enforced at query time by joining the policy to the dictionary, which means the policy engine now has a runtime dependency on application metadata.
  • Lineage is derived from the application's own field definitions rather than from parsing SQL, because the SQL refers to generic columns and carries no semantics.
  • The catalogue inherits a new staleness source. If the dictionary export lags by a day, newly created fields are unclassified and therefore, under a default-deny rule, invisible; under a default-allow rule they are exposed. Pick default-deny and accept the support load, because the alternative fails open.

What it cost them

Generality is paid for in query plans and in tooling. Nothing off the shelf understands the physical layout, so profiling, sampling, quality tests and masking all need the dictionary join, and each of those is a piece of platform work that a conventional warehouse gets for free.

When not to copy this

Below a few hundred tenants, do not copy it. Per-tenant schemas or a tenant column with row-level policies are cheaper, are understood by every tool you already own, and let you classify a column once for everyone. The metadata-driven shape earns its cost only when per-tenant DDL is the thing that would break, which means tens of thousands of tenants each customising their own fields. Adopting the architecture because a large vendor did is how a 200-tenant company acquires a catalogue problem it did not have.