Tenant Isolation in LLM Applications
Why a multi-tenant LLM application leaks through its caches, queues and memory rather than its database, how provider isolation units constrain your tenancy model, and what a shared prompt cache reveals through timing alone.
Two customers, one deployment. Every database row carries a tenant id, the retrieval collections are partitioned, and the access-control code has been reviewed twice. The incident, when it comes, will involve none of that. It will involve the semantic cache, the shared quota, the example pool someone mined from production traffic, or a trace viewer that a support engineer used to read another customer's prompt. The parts of an LLM application that hold tenant data are not the parts that look like storage.
Your tenancy model has to land on a provider isolation unit
Caches are the first place this bites, because a prompt cache is shared state by design and its sharing boundary is set by the provider, not by you. Anthropic's prompt caching is isolated per workspace on the Claude API, Claude Platform on AWS and Microsoft Foundry, while on Bedrock and Google Cloud the isolation is organisation-level only, and different organisations never share a cache under any circumstances (Anthropic, Prompt caching). A design built on "one workspace per tenant" is a real isolation primitive on the first set of platforms and a no-op on the second. The same unit governs money and throughput: spend limits are set per organisation and per workspace, and hitting one returns a 400 (Anthropic, Claude API errors), which makes the workspace the natural blast radius for a runaway tenant as well as for a cache.
Deciding this early is cheaper than retrofitting it, because the choice propagates into key management, cost attribution and every dashboard you build.
A shared cache is a side channel even when it returns nothing
The obvious cache risk is a wrong hit: a response conditioned on one tenant's data returned to another, which is the failure mode discussed in caching layers for LLM applications. The subtler risk is that a cache leaks through latency without ever serving a response across the boundary.
Gu and colleagues measured this directly. Because a cached prefix skips prefill, time to first token is data-dependent, and an attacker who can time requests can test whether a given prefix is already cached. Auditing seventeen API providers, they found global cache sharing across users in seven of them, which allows an attacker to learn something about other users' prompts, and they used the same timing signal to infer that OpenAI's embedding model is a decoder-only Transformer, a previously undisclosed architectural fact (Gu et al., 2025, Auditing Prompt Caching in Language Model APIs, ICML 2025, arXiv:2502.07776).
For an application built on top of a provider, the lesson transfers to every cache you run yourself. Isolating a cache per tenant removes the channel and costs hit rate: with a 2,000-token shared system prefix and cached reads priced at roughly a tenth of base input tokens, per-tenant isolation means paying one full-price cache write per tenant per TTL window instead of one for everyone. For a product with 50 active tenants and a five-minute window, that is 50 writes every five minutes rather than one, which is a real number you can put next to the risk rather than an argument you have to win on principle.
Noisy neighbours are a scheduling problem
Provider rate limits apply to your organisation, not to your tenants, so one customer's bulk job is capable of consuming the quota that every interactive user shares. The mitigation is admission control inside your own application: a per-tenant concurrency cap, a floor reservation for interactive traffic, and a fair queue that accounts in tokens rather than requests, because requests differ in cost by two orders of magnitude and a request-counting fair-share scheduler is not fair.
A workable split of a limit \(L\) tokens per minute across \(T\) tenants is to reserve a floor \(f = \alpha L / T\) per tenant and allocate the remaining \((1-\alpha)L\) by demand, with \(\alpha\) around 0.5 for a mixed workload. See admission control and load shedding for the mechanics and spend guardrails and quotas for the money side.
When it breaks
Traces are tenant data. A trace store that holds prompts and completions is a copy of everything your customers typed, usually with the weakest access controls in the system and the broadest internal audience. It needs the same tenant scoping as the primary store, plus redaction, plus an answer to how long it is kept.
Pooled training and memory are irreversible. An adapter fine-tuned on several tenants' data, or a shared long-term memory store, cannot be un-mixed; erasure means retraining. Keep per-tenant adapters, which multi-tenant serving and isolation shows how to serve cheaply, or keep the data out.
The agent is a confused deputy. One service identity holding the union of every tenant's tool credentials will eventually use the wrong one, because the model chooses the call and the credential is ambient. Scope credentials per request and let the tool layer reject a mismatch.
Residency moves under failover. Region selection is per request, so a fallback to another region relocates data. If a tenant has a residency commitment, the fallback policy has to be tenant-aware or absent.
Administrative limits arrive before technical ones. One workspace or project per tenant is clean until you reach the provider's cap on workspaces, keys or projects. Check the ceiling before the model depends on it.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Anthropic, Prompt caching platform.claude.com
- Anthropic, Claude API errors platform.claude.com
- Gu et al., 2025, Auditing Prompt Caching in Language Model APIs, ICML 2025, arXiv:2502.07776 arxiv.org
7 flashcards for this concept
Click a card to reveal the answer.