A reporting API offers an export endpoint. For most customers it returns in eight seconds. For the largest it must process about 40 million rows, takes six to nine minutes, and the load balancer kills the connection at 60 seconds. Clients retry, which starts the export again. Design the endpoint.
Show the full answer Hide the answer
Requirements that drive the structure
Three facts settle the shape before any technology is chosen. The work takes minutes, not seconds, so it cannot live inside a request. The duration is data-dependent, so there is no timeout that is both safe for the largest customer and honest for the smallest. And the client retries, so a request that starts work must be safe to repeat.
A synchronous endpoint cannot satisfy any of these. Raising the proxy timeout to ten minutes trades one failure for a worse one: connections held open through every hop, a thread or socket per in-flight export, and retries that multiply the work because the client cannot tell a dropped response from a dropped request.
The design
POST /exports accepts the request, starts nothing synchronously, and returns 202 Accepted with a Location header pointing at an operation resource. The operation resource is the contract:
GET /exports/9f3c → 200
{ "status": "running", "created_at": "...", "percent_complete": 46,
"retry_after_seconds": 15, "result": null, "error": null }
- Terminal states are explicit and final:
succeeded,failed,cancelled. A client that sees a terminal state stops polling and never has to guess. Retry-Afteron the poll response carries the server's pacing, so a thousand clients do not choose their own interval. Without it the poll traffic becomes its own load problem: 40,000 subscribers polling every second is 40k rps of pure overhead.- The result is a separate, expiring artefact, not an inline payload. A signed URL to object storage with a 24-hour lifetime keeps a 3 GB file out of your API tier entirely.
- The creation call takes an idempotency key. The same key returns the same operation id instead of starting a second export. This is what makes the client's retry safe, and it is the part most designs omit.
- Cancellation is a
DELETEon the operation, which matters because a customer who mis-filtered a 40-million-row export should not have to wait for it.
What it costs
You now run a job system: a queue, workers, state in a durable store, a reaper for operations nobody collected, and quotas on concurrent exports per tenant. The client is more complex too, because polling logic and terminal-state handling replace one blocking call. Expect a week of work and a permanent operational surface rather than an afternoon.
The second cost is conceptual. The operation resource is a public contract: its states, its id format and its result lifetime are now things integrators depend on, and shortening the result lifetime later will break someone. Publish the retention period as part of the API, not as an implementation note.
How it fails
The two common failures are both bookkeeping. An operation stuck in running forever because the worker died without updating state, which a lease with a heartbeat and a timeout-to-failed transition fixes. And result URLs that outlive their authorisation, where a signed link to a report full of personal data is still live weeks later in somebody's chat history. Scope the signature to a short window and to the requesting principal.
When this is over-engineering
If the slowest case is four seconds and bounded by a hard row limit, keep it synchronous and cap the request. A marketplace data feed that nobody waits on in real time can be simpler still: generate on a schedule, publish to a known location, and let clients fetch the latest file. That removes the operation resource, the queue and the polling. Reach for async request-reply when duration is genuinely unbounded by the customer's data, not because the endpoint feels slow.