Ticketmaster stated that the November 2022 Eras Tour onsale drove 3.5 billion total system requests, about 4x its previous peak, attributing much of it to bot traffic and users without invite codes. At 4x peak with inventory far below demand, you can queue, fail fast, or degrade. What does each buy, and what does it cost?
Show the full answer Hide the answer
What is gained, and in which currency
The three options optimise different things, and the mistake is treating them as points on one scale.
Queue. Admit a bounded number of users to the purchase path and hold the rest in an ordered waiting room. Buys fairness and a bounded load on the inventory system, which is the component that cannot be scaled by adding instances because it is serialising on contested rows. The queue is the only option that lets the backend run at its designed capacity rather than above it.
Fail fast. Reject above a measured concurrency limit with a fast error. Buys survival of the fleet, cheaply and immediately. A fast rejection costs a connection and a few milliseconds; a slow failure holds a thread for 30 seconds and takes capacity with it.
Degrade. Keep the purchase path and shed everything else: no recommendations, no seat maps, no personalisation, cached event pages, read-only browse. Buys the most transactions completed per unit of capacity, because it reallocates capacity from the 95% of requests that are not buying to the 5% that are.
What is paid
- Queueing pays in truth about the wait, and in abandonment. A queue that cannot say "you are position 180,000 and inventory will be gone before you arrive" is a mechanism for making people wait for nothing — and when demand exceeds supply by orders of magnitude, that is most of them. The queue must be able to close, telling later arrivals the truth rather than enqueuing them. It also needs its own scaling, independent of the thing it protects, and the waiting room becomes the new target for the bots.
- Failing fast pays in randomness. Who gets through is decided by retry timing and network luck, which is the least defensible allocation when the goods are scarce and the event is public. It also invites client retry storms, since the natural response to a fast error is an immediate retry, so it requires enforced backoff to work at all.
- Degrading pays in the features it removes, and some of those features are load-bearing. On a ticketing flow the seat map is close to the product, so shedding it reduces completed purchases rather than protecting them. Degradation also has to be built in advance — a path that serves without recommendations has to exist and be tested — which is why it is unavailable to teams that need it for the first time during the event.
When the bill arrives
All three bills arrive in the first ten minutes, and at 4x peak the distinguishing factor is which component is the binding constraint. If the constraint is the inventory system's write contention, only the queue helps, because fail-fast and degradation both still admit more concurrent purchase attempts than the rows can serialise.
The second bill is reputational and arrives later: the architecture visibly allocates a scarce good, so the mechanism becomes a public fairness question, which is a constraint no amount of capacity planning addresses. The stated 4x figure and the bot attribution are an argument about demand, and demand far above inventory is not a load-shedding problem at all.
The answer, and how to keep the option to reverse
All three, layered, with the queue as the admission mechanism: verified-identity admission before the onsale to make demand countable, a queue sized to actual inventory with honest position and a closed state, concurrency limits inside the queue as the fail-fast backstop, and a pre-tested degradation ladder shedding non-purchase features in a defined order.
Keep each layer independently switchable with a flag, and rehearse the ladder. A degradation path that has never been exercised does not exist, and the onsale is not the moment to discover which features were load-bearing.
When not to build this
Below roughly 10x your normal peak, none of it is justified. Concurrency limits with fast rejection, plus autoscaling with real headroom, plus a static cached page for browse, handles an ordinary marketing spike. The queue in particular is a distributed system with its own failure modes and its own on-call burden, and buying it for a 2x Black Friday is buying an outage source to prevent a latency rise. The threshold is a demand-to-inventory ratio where fairness becomes a public question, not a traffic multiple on its own.