Review this documentation. The reference says "the API is limited to 100 requests per second" and lists 429 in the error table. Integrators still get throttled constantly, build their own guesswork pacing, and open tickets asking what the real limit is. What is missing, and what would you publish instead?
Show the full answer Hide the answer
What is actually required
A client cannot pace itself against a number. It needs to answer four questions in code: am I close to the limit, what happens when I cross it, when may I try again, and does this request cost more than another. "100 requests per second" answers none of them, which is why integrators reverse-engineer the limiter by hitting it in production.
What I would remove, and what is safe to remove
Delete the bare number. It is worse than nothing because it reads as a promise, and the first time a client sustains 95 rps and still gets throttled, your documentation has lied to them. Delete the 429 row from the generic error table too, and replace it with a section, because throttling is a protocol a client implements rather than an error it logs.
The one change that matters
Publish the limiter's shape, not its headline number:
- Which key the limit is counted against — the API key, the account, the account plus endpoint group, or the IP. Integrators running one key across twelve workers need this to decide whether to coordinate.
- Sustained rate and burst separately. A token bucket of 100 per second with a 500-token bucket behaves completely differently from a fixed 100-per-second window, and only one of them tolerates a six-second batch. State the algorithm; it is not a secret and a client that knows it can pace perfectly.
- The cost model where requests are not equal. A search returning 500 results or an inference call with a 100k-token prompt is not one unit of anything. If the limit is really in units of work, publish the units and the per-endpoint cost, as expensive APIs increasingly do.
- The exact 429 contract: that
Retry-Afteris always present, whether it is seconds or a date, and that honouring it is required. A documented 429 without a documentedRetry-Afterguarantees every client invents its own backoff, and the aggressive ones become your next incident. - Live budget headers. Returning the remaining quota and the window reset on every response lets a client slow down before it is rejected, which is cheaper for both sides than rejection plus retry. The IETF's
RateLimitandRateLimit-Policyfields are the emerging convention; as of late 2026 they are still an Internet-Draft, so publish your header names explicitly rather than assuming a client knows them. - What happens at the limit beyond 429: whether sustained abuse escalates to a longer block, and whether limits differ by plan. Surprise is the thing integrators are complaining about.
What I would leave alone
Keep the number in the plan comparison for sales, where an order of magnitude is the useful content. And keep the limits conservative in the free tier: the goal of the documentation change is not a higher limit, it is a client that can be built correctly against the limit you have.
How I would argue this in review
Measure it. Count 429s per integrator and the share of integrators whose 429 rate is non-zero in a normal week. If a third of your integrators are being throttled routinely, they are all running retry loops against you, and every one of those loops consumes the same quota as a first attempt. The documentation change is the cheapest available reduction in your own load, and it is a week of writing rather than a quarter of engineering.
When not to publish all of this
A single-tenant internal API whose only caller is one of your own services does not need a published limiter contract; the two teams can agree a number in a channel and move on. The documentation above earns its keep once callers outnumber the people you can talk to, and the threshold is roughly the point where you stop knowing every integrator's name.