Cell Affinity Leak
also called Cell Identifier Escape, Non-Movable Tenant
The condition where a cell's identity has escaped into client-visible artefacts - hostnames, webhook URLs, allow-listed addresses, exported ids - so a tenant can no longer be moved between cells without breaking its integrations.
A software-as-a-service platform runs 12 cells, each a complete stack serving a subset of tenants. One tenant has grown to consume 40% of its cell, so the plan is the obvious one: move it to a cell of its own.
The move turns out to be impossible this quarter. The tenant's API hostname is cell7.api.example.com, which
appears in their integration code. Their webhook callbacks were registered from cell 7's egress addresses,
which their security team allow-listed in a firewall change that took six weeks. Exported report files carry
ids with the cell number in them, and their data warehouse joins on those ids.
Nothing about the cell architecture was built wrong. The cell's identity simply leaked out of the infrastructure and into other companies' configuration, and a tenant whose cell cannot change is a tenant whose cell is a capacity ceiling rather than a blast-radius boundary.
Why it matters
Cells deliver two things: a bounded failure domain, and the ability to rebalance load by moving tenants. The second is what pays for the first. A cell you cannot rebalance becomes permanently hot, and the response is to over-provision every cell to the size of its largest possible tenant, which removes most of the economic argument for cells in the first place.
The leak is also discovered at the worst time. Nobody tries to move a tenant until a cell is already struggling, so the constraint surfaces during a capacity incident, and the only remaining options are vertical growth of one cell or a migration negotiated with the customer over months.
Implementation patterns
- One stable hostname for everybody, with cell selection at the edge. Per-cell hostnames are the most common leak and the hardest to withdraw.
- Cell identity resolved through a router keyed by tenant, with the mapping held in a control plane — the opposite of embedding the location in the identifier, which is a fine design for shards you never move and the wrong one for cells you must.
- Cell as a short-lived token claim, re-issued by the router, never a constant the client stores.
- A written movability contract: the list of artefacts that may never contain a cell reference — identifiers, URLs, filenames, error messages, support ticket fields — enforced by a test over the public API surface rather than by review.
- Shared egress addresses across cells, so a customer's allow-list survives a move.
- A rehearsed move with a measured duration. The planning number is how long a median tenant takes: 2 TB copied at 200 MB/s is roughly 3 hours plus cut-over. If that number is unknown, rebalancing is theoretical.
Industry example
Any multi-tenant platform in production that started with one cell per region and gave each a regional hostname has some version of this. The pattern that survives is a single global endpoint with routing behind it, and the pattern that traps teams is the one that looked more transparent: telling enterprise customers which cell they are on, which they then encode. Client-side DNS caching makes the hostname case worse than it looks — customers cache far longer than any TTL you publish, so even withdrawing a per-cell name takes quarters.
Failure scenarios
- A hot cell that cannot be relieved, so one tenant's growth degrades everyone in the cell.
- Tenant moves that break webhooks silently, because the callback source address changed and the customer's firewall drops it with no error visible to you.
- Cell-encoded ids in a customer's warehouse, making the move a data migration in someone else's estate.
- An emergency evacuation that cannot run when a cell is damaged, which is the scenario cells exist for.
- Support tooling keyed on cell rather than tenant, so the move breaks internal processes too.
Trade-offs
Opaque routing costs a lookup on every request and a control plane that must be more available than the cells it fronts, because every request depends on it. Per-cell hostnames are cheaper, simpler to debug and nicer for customers who want to point a firewall rule at something. The leak is cheap to prevent on day one and expensive to repair later, which makes it a decision to take before the first external integration, not after.
When not to use it
If the product genuinely has one cell and will never have more — a single-region appliance, an internal platform with a fixed tenant set — the indirection buys nothing and a direct hostname is clearer. The threshold is the first external integration: once another company's configuration can contain your topology, the cost of the leak stops being yours to control.
Interview question
Q: You have eight cells and one tenant is now 40% of its cell. Walk me through what you check before committing to a move date, and what you would change in the platform so the next move is routine.
What a strong answer covers: inventory of every place the cell id is observable to the customer · webhook egress and allow-lists · exported identifiers and file names · the measured copy time for that tenant's data and the cut-over window · the platform changes that make moves routine (single endpoint, tenant-keyed routing, shared egress, opaque ids) · and a rehearsal cadence, because a move path that has never been exercised is not a capability.
Quick check
Quiz: Why does giving each cell its own hostname undermine cell-based architecture? It leaks cell identity into customer configuration, so tenants cannot be moved or evacuated, and the cell stops being a rebalanceable unit.
Flashcard: What single property makes a cell architecture worth its cost, and which design detail usually removes it? — The ability to move a tenant between cells; per-cell hostnames, allow-listed per-cell egress addresses and cell-encoded ids remove it by putting the topology in other companies' configuration.