A signup flow returns HTTP 202 and shows "code sent" as soon as it has handed the SMS to a provider. Twilio documents that traffic above a sender's throughput is queued rather than rejected, that a sender's queue holds at most four hours of messages before they fail (error 21611 on a single number, 30001 through a Messaging Service), and that a US long code sends roughly one message segment per second. A marketing push multiplies signups 14x for two hours. What did the 202 buy, and when does the bill arrive?
Show the full answer Hide the answer
What is gained
The request path stops waiting on a carrier hop that is tens of milliseconds at p50 and seconds at p99, plus any provider-side retry. Signup p99 becomes the application's own work, typically under 100 ms. More importantly, the caller's availability stops compounding with the provider's: a provider blip no longer fails signups, it delays codes. That is a real and correct gain, and it is why the async boundary belongs here.
What is paid
The 202 changed what the acknowledgement means. It is now a promise about durability, not about delivery - the message is accepted somewhere and will be attempted. The screen says "code sent", which is a stronger claim than the 202 supports, and the gap between the two is where the incident lives.
The arithmetic: a long code at roughly one segment per second delivers about 3,600 messages an hour. A 200-per-hour baseline is comfortable; at 14x it is 2,800 an hour, still under the ceiling - until you add resends from users who did not get a code, and any message long enough to bill as two segments. Cross the sender rate and the backlog grows, and the queue is bounded in time, not in messages: four hours of that sender's own rate. Past that, messages fail rather than wait.
When the bill arrives
Not at the moment of saturation. The queue absorbs the first few minutes, so dashboards look fine and the only symptom is support tickets. By the time depth is visible, the codes being delivered were requested 20 minutes ago and are useless - a one-time code with a 10-minute validity that arrives after 20 minutes is a failed signup that the system recorded as a success. The async boundary did not create downstream capacity. It converted rejection into latency, and the bounded buffer converts latency back into rejection at the one place the caller decided not to look.
What to change
- Set a validity period close to the code's own usefulness. Twilio lets you shorten the queue window per message or per Messaging Service; a 10-minute code should carry roughly a 10-minute validity so the provider fails fast instead of delivering something already expired.
- Raise aggregate throughput rather than the queue. A Messaging Service spreads traffic across senders; a short code moves the ceiling by two orders of magnitude relative to a long code.
- Alert on queue age, not queue depth. Depth tells you how much work there is; age tells you whether the promise is still true.
- Make the UI say what the 202 means - "sending", with the state updated from the delivery callback and a visible retry - so the user's model matches the system's.
When the synchronous call was the right answer
When the user is sitting there waiting for exactly this outcome and the downstream is fast and rarely down. A 2-second synchronous attempt that fails visibly beats a 202 that may quietly expire in four hours, because the user can act on an error and cannot act on a false success. The general rule: go async when the caller's outcome does not depend on the result, and when it does, go async only if you are willing to build the state machine, the status surface and the expiry that make the eventual result visible. Teams buy the async boundary and skip the three things that make it honest, which is how a latency improvement becomes a silent failure mode.
Common weak answers
- "Add retries." The provider already retries. Client retries on a saturated sender add segments to a queue that is failing on age.
- "Increase the queue size." It is time-bounded by design, and a deeper queue delivers more codes nobody can use.