In March 2023 OpenAI disclosed that a bug in the redis-py async client let one request receive data left behind in a recycled connection, so some users saw other users' chat titles and some payment details. The authorisation code was not at fault. Which class of control was missing, and what would you change?
Show the full answer Hide the answer
The situation they were in
A Python server using asyncio talks to a Redis cluster through a shared connection pool, which is the normal and correct design: connections are expensive, so they are borrowed, used and returned. On 20 March 2023 a server-side change caused a spike in Redis request cancellations.
The mechanism
In the async client, a request and its response behave as two queues on one connection. The caller pushes the request, then pops the response, then returns the connection to the pool. If the caller is cancelled after pushing the request but before popping the response, the connection goes back to the pool with an unread response still in it. The next borrower pops that response and believes it is its own.
The race existed before the change and was rare. The cancellation spike raised its probability enough to be visible to users, which is the general shape of this class of incident: a latent correctness bug in shared infrastructure, made frequent by an unrelated load change.
What it cost them
Conversation titles from the history sidebar were visible to other users, and in some cases the first message of a newly created conversation. OpenAI reported that payment-related information for 1.2% of ChatGPT Plus subscribers active during a roughly nine-hour window on 20 March may have been visible to another active subscriber: name, email address, payment address, and the last four digits and expiry date of a card. The service was taken offline to fix it, and the postmortem was published within days.
The control that was missing
Nothing between the cache and the response re-checked that the data belonged to the principal who asked for it. Authorisation happened at the front door. Past it, bytes arrived from shared mutable infrastructure that had no concept of a user.
The general control is request-scoped identity at the point of use: store the owning subject alongside any cached or pooled value, and compare it to the request's principal before serialising. The cost is one comparison and a few bytes per entry. The second control is at the pool: a connection whose protocol state machine is not known-clean should be discarded rather than returned. That is strictly safer and it churns connections exactly when cancellations spike, which is why libraries are reluctant to do it by default.
Detection deserves equal weight. This bug produces no errors and no latency change. The only signal is the owner comparison above, alerting on a non-zero daily count, or a user telling you.
Where copying the fix would be a mistake and when not to generalise it
Do not conclude that shared pools are the problem. Pooling is why the system is fast, and one connection per request is not viable at that volume. Equally, do not add an owner check to every object on every path, which buys little and costs clarity. Add it where shared mutable state sits between authorisation and the response, which you find by asking one question of a design: what is reused across requests here, and is it keyed by principal or merely keyed by something that usually differs? Caches, connection pools, thread-local context, template contexts and prepared-statement handles are the usual list.