Tenant-Scoped Layout Key
also called Tenant Clustering Key, Tenant-Aligned Physical Layout
Making the tenant identifier the column that governs a table's physical organisation, so isolation, per-tenant cost attribution and per-tenant deletion come from layout rather than from access control alone.
A shared analytics platform holds 40 TB of event data for 900 internal tenants. Access control is
correct: row-level policies work, nobody sees another tenant's data. Three problems remain, and
none of them are access problems. One tenant's query scans the whole table and slows everyone.
Nobody can say what any tenant costs. And a departing tenant's erasure request means a DELETE
that rewrites terabytes and leaves the old files alive in snapshot history.
A tenant-scoped layout key is the decision to make tenant_id the column that determines
physical organisation - the partition key, the clustering key or the sort key - so that these
three become properties of the storage rather than problems solved on top of it.
Why it matters
Access control answers "may this principal read this row". It says nothing about how many bytes were read to find out, whose bytes they were, or how to remove them later. Layout answers all three at once, because a query filtered on tenant reads only that tenant's files, the engine's own statistics report bytes scanned per partition, and removal becomes a file operation.
The scan effect is the one people underestimate. Without tenant-aligned layout, every file's
tenant_id range spans the whole key space, so no file can be skipped and a query for one small
tenant reads the same volume as a query for the largest one. With it, the small tenant's query
reads a proportional slice, which changes both latency and cost by an order of magnitude on a wide
table.
Implementation patterns
- Cluster or sort on tenant rather than partitioning on it, unless tenant count is low and volumes are even. Partitioning on a high-cardinality tenant column produces thousands of tiny partitions and a metadata problem that costs more than the pruning saves.
- Compound the layout: partition by the retention-bearing date column, sort by tenant within the partition. That keeps retention and time predicates working while making tenant filters selective.
- Separate the whale tenants. When the largest handful hold most of the volume, give them their own tables or their own partitions and leave the long tail sharing, because one layout cannot serve both distributions.
- Make tenant derivable from the key where you control key generation. Pinterest described in 2015 packing a shard number into the high bits of a 64-bit object id so routing needed no lookup; the analytical version of the same idea is that tenancy recoverable from the key survives joins and does not need carrying as a separate column.
- Attribute cost from partition statistics, not from a tagging scheme that has to be maintained.
- Test erasure as a drill. Time an actual tenant removal, including snapshot expiry, and compare it to the contractual deadline.
Industry example
Pinterest's 2015 description of its sharded MySQL fleet is the clearest public statement of the underlying idea, in a transactional setting: a 64-bit id composed of a 16-bit shard number, a 10-bit type and a 36-bit local sequence, so any object routes to its host arithmetically. The transferable part is not the bit layout, which solves a routing problem analytics does not have. It is that an identifier scheme is the hardest thing in a platform to change, so whatever the layout will be organised around belongs in the key from the beginning.
Failure scenarios
- Skew. Five tenants holding 80% of the data become five partitions that dwarf the other 895, and the largest tenant's queries are as slow as they ever were while everyone else's improved.
- Partition explosion. Partitioning directly on a high-cardinality tenant column turns one day's 300 MB into hundreds of files of a few megabytes, and planning cost overtakes scan cost.
- A key that cannot be changed. Two tenants merge after an acquisition and every row of both must be rewritten because tenancy was inside the primary key rather than beside it.
- Layout mistaken for security. Tenant-aligned files make a query cheap; they do not stop a principal with table-level access from reading everything.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Partition by tenant | Erasure is a drop; perfect isolation of scans | Severe skew and partition explosion above a few hundred tenants |
| Cluster or sort by tenant | Most of the pruning; no metadata explosion; reversible | Erasure is still a rewrite; benefit degrades between maintenance runs |
| Tenant inside the surrogate key | Pruning survives joins; no extra column | The key becomes permanent; merges and splits are full rewrites |
When not to use it
When most queries are cross-tenant, the layout key is spent on the wrong column. A platform whose main workload is aggregate analysis across all tenants gains nothing from tenant-aligned files and loses the ordering that would have made its real predicate cheap, since a table has one physical order. The same applies when tenant volumes are tiny and even - a few hundred megabytes each means every query is cheap regardless - and when the isolation requirement is contractual rather than performance-related, in which case separate tables or separate databases per tenant are the honest answer and layout is a distraction.
Interview question
Q: A shared analytics platform has 900 tenants, correct row-level security, and complaints about noisy neighbours and unattributable cost. An engineer proposes putting tenant into the primary key. What do you ask before agreeing, and what would you propose instead?
What a strong answer covers: asking for the tenant volume distribution first, because skew
decides everything that follows; asking whether tenants are ever merged or split, which is what
makes a key irreversible; distinguishing the three problems - isolation, attribution and erasure -
because they have different cheapest solutions; proposing clustering on an ordinary tenant_id
column as the reversible move that captures most of the benefit; and naming the measurement that
would justify going further, which is bytes scanned per query for a small tenant against the table
size.
Quick check
Quiz: Row-level security is correctly configured on a shared analytics table. Which two tenancy problems does it still not solve? - Cost attribution and scan isolation. A policy decides which rows are returned, not how many bytes were read to find them or whose workload paid for it.
Flashcard: Why is a high-cardinality tenant column usually a clustering key rather than a partition key? - Because partitioning on it multiplies partition count and divides file size, so per-file planning cost overtakes the pruning benefit, while clustering gives file skipping without the metadata explosion.