A browser-based design tool in Canva's mould lets users export a 40-page document to PDF, a 90-second animation to MP4 or a print-ready file. Export work takes anywhere from 3 seconds to 11 minutes, about 60000 a day, bursting around 09:00 in each region. All three run inside the synchronous web tier today and p99 export requests are killed by the load balancer at 60 seconds. Which architecture style should the export path use?
Show the full answer Hide the answer
The deciding property
One fact settles it, and it is not the volume. The duration distribution spans more than two orders of magnitude, from 3 seconds to 11 minutes. No request/response style can hold a connection across that range: the client, the load balancer, the gateway and the browser tab each impose their own ceiling, and the work does not care about any of them. Once a unit of work can outlive the request that asked for it, the architecture owes the caller a job identifier and a way to ask about it later, and everything else follows from that.
60000 a day is about 0.7 per second averaged, which sounds trivial and is exactly why the volume is the wrong number to reason from. The bursts and the long tail are what matter: a regional 09:00 burst of long print exports can occupy every worker while three-second PDF jobs queue behind them.
The design the answer implies
Submit returns 202 with a job id. The job goes on a durable queue, partitioned by output type so a print job cannot starve a PDF job. A worker pool per type is sized independently, because a PDF worker needs memory and an MP4 worker needs CPU and neither should be scaled by the other's load. Progress and completion reach the client by polling the job resource or over an existing socket. Output lands in object storage with a signed URL.
The durability is not optional. With 60000 jobs a day, a deploy that discards in-flight work loses hundreds of user-visible exports, and the users who notice are the ones who waited eleven minutes.
Why the other options fail
One synchronous microservice per output type is the most popular wrong answer because it looks like progress. It changes where the work runs and not the shape of the interaction: the gateway still holds a connection for 11 minutes, the 60-second ceiling still kills the long tail, and a rolling deploy still destroys in-flight renders. It also buys three deployables, three pipelines and three on-call surfaces for no property the current design lacks.
A space-based architecture with an in-memory data grid answers a different question. That style exists for extreme read/write contention over shared state with a database that cannot keep up. Export rendering is embarrassingly parallel and touches a document per job with no contention at all. The grid adds a distributed cache to operate and solves nothing.
Raising the load balancer timeout to 15 minutes treats the symptom and creates three new problems: a browser tab must stay open for the whole render, every deploy kills in-flight exports, and 15-minute connections make connection-slot exhaustion the next incident. It also makes the timeout a load-bearing configuration value that no one will dare change.
What would flip the choice
| If this changes | Choose | Because |
|---|---|---|
| Every export finishes under 2 seconds | In-process with a bounded thread pool | A queue's operational cost buys nothing when the work fits inside the request |
| Volume rises to millions a day with SLA tiers | Queue plus priority classes and admission control | Fairness across tenants becomes the hard problem rather than duration |
| Renders need exactly-once billing | Queue plus an idempotency key per job | At-least-once delivery will replay a job and double-bill without one |
When this is the wrong answer
For a team of four shipping a first version, a database table used as a queue and a single worker process is the right answer and a message broker is not. It handles thousands of jobs a day, it is inspectable with SQL, and it needs no new infrastructure. Move to a broker when you need per-type isolation, delayed retries and consumer scaling that the table is making painful — and not before you have felt that pain.