pattern

Request-Scoped Identity

also called Principal-Bound Response, Per-Request Ownership Check

Binding every piece of shared mutable state a response is assembled from to the principal who asked for it, so that a bug in pooling or caching produces an error rather than another customer's data.

openaimulti-tenancyconnection-poolcachingauthorization

Authorisation almost always happens at the front door. A gateway establishes the principal, checks a scope, and hands the request onward. Everything past that point assembles a response from caches, connection pools, async context and prepared statements, none of which have any concept of a user.

That gap is where cross-tenant leaks come from. The authorisation code is usually correct. What fails is a shared resource reused across requests that returned the wrong occupant's data: a cache keyed on a resource id but not the tenant, a connection returned to the pool with an unread response in it, a context mutated by a previous request. Each produces a response that passes every check and belongs to somebody else.

Request-scoped identity makes ownership travel with the data rather than with the request. Every cached value and every pooled resource is either keyed by principal or carries the principal with it, and the comparison happens where the value is used.

Why it matters

This class of leak emits no error, no latency change and no log line, so the first signal is a customer reporting that they saw somebody else's data, and the exposure window is however long the bug ran.

The economics are lopsided. Storing the owning subject with a cached value costs a few bytes and one comparison, nanoseconds against a lookup already costing microseconds. The alternative is a breach notification.

Implementation patterns

  • Principal in the cache key. order:4821 is wrong; tenant:17:order:4821 cannot be served to tenant 18 even if the application asks for it.
  • Owner stored with the value and asserted before serialisation, which also catches a key computed wrongly. The two compose and the second is stronger.
  • Authorisation in the data-access layer, so a repository requiring a principal argument makes an endpoint that forgets to ask return nothing rather than everything.
  • Discard rather than recycle on uncertainty. A pooled connection whose protocol state is not known-clean, typically after a cancellation or timeout, is closed instead of returned. It costs reconnects exactly when the system is stressed, which is why libraries default the other way.
  • An assertion metric: count owner mismatches and page on any non-zero daily total. For a failure mode that emits nothing else, this is the only detector you get.

Industry example

In March 2023 OpenAI disclosed an incident in ChatGPT caused by a bug in the asynchronous redis-py client. A request cancelled after pushing its Redis command but before reading the reply returned the connection to the pool with that reply still queued, and the next borrower read it. A server-side change on 20 March raised the cancellation rate and made a rare race frequent.

Users saw other users' conversation titles and in some cases the first message of a new conversation. OpenAI reported that payment-related information for 1.2% of ChatGPT Plus subscribers active during a roughly nine-hour window may have been visible: name, email, payment address, and the last four digits and expiry of a card. The authorisation layer was not at fault. Nothing between the pool and the response re-checked ownership.

Failure scenarios

  • A cache keyed on resource id alone, where two tenants legitimately share an id space.
  • A connection pool corrupted by cancellation, serving a previous caller's reply.
  • Async or thread-local principal not cleared, so a pooled worker retains the previous principal when a request fails before setting it.

Trade-offs

Choose Gains Pays
Principal in every cache key Cross-tenant hits become impossible Lower hit rate for genuinely shared data
Owner stored with the value Catches wrong keys as well as wrong lookups A few bytes per entry and a hot-path comparison

When not to use it

In a single-tenant system where every reader is entitled to every record, principal-keyed caches cost hit rate and buy nothing. The same holds for public data: a catalogue page should be cached once, not once per visitor, and keying it by principal can multiply origin load by the number of active users.

Apply it where a shared mutable resource sits between an authorisation decision and a response, and where its readers are not all entitled to the same data. Applying it everywhere turns a security invariant into a performance problem and trains people to remove it.

Interview question

Q: "Your API authorises correctly at the gateway and a customer still reports seeing another customer's record. The authorisation code is unchanged and error rates are flat. Where do you look, in what order, and what single invariant would you write into the codebase to stop the class rather than the instance?"

What a strong answer covers: the shared-state inventory before the authorisation code; flat error rates as evidence for this class rather than against a bug; the invariant stated as "every cached or pooled value carries its owning principal, asserted at the point of use"; and the limit, that public data should stay shared.

Quick check

Quiz: A cache key is order:{id} and two tenants have overlapping order ids. Where is the leak and what is the fix? Answer: on any cache hit for a colliding id; put the tenant in the key and store the owner with the value.

Flashcard: Why does a cross-tenant leak show flat error rates? The failure is a correct response built from the wrong data, so nothing throws.