Operation Resource
also called Async Request-Reply, Long-Running Operation, Job Resource
A separately addressable resource representing work that outlives its request, so a client can submit once and then poll for status and results instead of holding a connection open.
An export endpoint returns in eight seconds for most customers. For the largest it processes about 40 million rows, takes six to nine minutes, and the load balancer kills the connection at 60 seconds. The client retries, which starts the export again. The work is minutes long, its duration depends on the customer's data, and the client cannot tell a dropped response from a dropped request — three facts that no synchronous design satisfies.
An operation resource splits submission from completion. POST /exports returns 202 Accepted with a Location header naming a new resource, and that resource is the contract: its status, its terminal states, its poll pacing and the lifetime of its result.
Why it matters
Raising a proxy timeout trades one failure for a worse one. Long-held connections occupy a socket and often a thread at every hop, they fail in the middle for reasons unrelated to the work, and the client's only recovery is a retry that duplicates minutes of compute. Capacity then scales with concurrent waiting rather than with work done.
Making the operation addressable also makes it operable. You can cancel it, report progress, retry it server-side, expire its output, and quota it per tenant, none of which is expressible when the job exists only as a stack frame inside a request.
Implementation patterns
- Terminal states are explicit and final:
succeeded,failed,cancelled. A client that sees one stops polling and never guesses. - The server sets the poll interval with
Retry-Afteron the status response. Without it, clients choose their own, and 40,000 subscribers polling once a second is 40k rps of pure overhead before any work happens. - Creation takes an idempotency key, so the same key returns the same operation id instead of starting a second job. This is the part most designs omit, and it is what makes the client's retry safe.
- The result is a separate expiring artefact, typically a signed object-storage URL with a published lifetime, which keeps a multi-gigabyte payload out of the API tier.
DELETEon the operation cancels it, which matters when a customer mis-filters a 40-million-row export.- A lease with a heartbeat, so a worker that dies moves the operation to
failedrather than leaving itrunningforever. - Per-tenant concurrency quotas, because the natural abuse of a cheap
POSTis a hundred of them.
Industry example
The shape is conventional across large platforms: cloud providers' long-running-operation conventions (documented in Google's public API design guidance and in Azure's REST guidelines), bulk and feed APIs on large marketplaces, and report-generation APIs in analytics products all return an operation or job handle and expect the client to poll. It is also what a well-built batch export looks like when the vendor does not call it an API: a marketplace data feed generated on a schedule and published to a known location is the same pattern with the operation resource replaced by a filename, and for consumers who do not need it on demand that is the cheaper design.
Failure scenarios
- Operations stuck in
runningbecause a worker died silently. The symptom is a slowly growing count of non-terminal operations older than the p99 duration, which is the alert to build. - Result URLs that outlive their authorisation, leaving a signed link to a report full of personal data live in somebody's chat history for weeks.
- Poll storms when
Retry-Afteris absent and clients implement a one-second loop, so the status endpoint becomes the hottest path in the system. - Retries without an idempotency key producing three identical exports and triple the compute bill for one customer.
- An operation store that is not durable. Keeping state in memory means a deploy loses every in-flight job and clients see operations that have ceased to exist.
Trade-offs
What you gain is a request path whose latency no longer depends on the customer's data size, and work that survives connection failures and deploys. What you pay is a job system: a queue, workers, durable state, a reaper, quotas, and client code that handles polling and terminal states. Expect a week of engineering and a permanent operational surface rather than an afternoon. The second cost is contractual: operation ids, state names and result lifetimes are now public, and shortening the retention later will break an integrator.
When not to use it
If the slowest case is a few seconds and bounded by a hard row limit, keep it synchronous and cap the request — the limit is the simpler mechanism and it is self-documenting. If consumers want the data on a schedule rather than on demand, a published file drop removes the operation resource, the queue and the polling entirely. And if the work is short but bursty, a queue in front of a synchronous endpoint with a 503 and Retry-After is less machinery than an operation resource for the same protection.
Interview question
Q: You are adding a bulk import endpoint to an API used by 2,000 integrators. Walk me through the resource design, and tell me specifically what happens when a client submits the same import twice because its first attempt timed out at the load balancer.
What a strong answer covers: 202 plus Location, terminal states, server-controlled polling, and the result as an expiring artefact · the idempotency key as the answer to the duplicate-submission question, including what scope and retention the key record needs · partial-failure semantics for an import, where row-level errors must be retrievable without re-running · per-tenant quotas and cancellation · and the explicit statement that the operation's id format and result lifetime are public contract.
Quick check
Quiz: Which single field on the status response most reduces the load an operation resource creates, and why? — Retry-After, because it moves the poll interval from the client's guess to the server's budget, and polling overhead otherwise scales with subscriber count rather than with work.
Flashcard: Why does a long-running operation endpoint need an idempotency key on creation? — Because the client cannot distinguish a lost request from a lost response, so its retry would otherwise start a second job and double the compute.