intermediate 3 min answer Multiple choice

Checkout with a p99 target of 150 ms and image processing that tolerates minutes run inside one deployable sharing a thread pool. Image bursts have saturated CPU and caused three checkout incidents this quarter. What is the smallest change that fixes it?

architecture stylesisolationbulkheadsdeploymentworkload
Pick one
Show the full answer Hide the answer

The deciding property

The two workloads need different units of failure and different units of scaling. Nothing in the situation says they need different code, different data or different teams. That distinction settles the question, because isolation is produced by the resource boundary - process, pool, node, queue - and not by the repository boundary.

The style question here is not "monolith or microservices". It is "which resource is shared", and the answer today is the worst possible one: a thread pool and a CPU allocation.

Why the second deployment

Deploying the same artefact twice, with the image workers consuming a queue and checkout serving HTTP, gives separate CPU and memory limits, separate scaling signals (queue depth for one, request rate for the other), separate restarts and a separate saturation domain. Checkout latency stops depending on image queue depth, which was the actual defect, and the work is hours of pipeline configuration rather than a migration.

What you keep is real and worth stating: one codebase, one database, one release. A deploy still ships both, and a memory leak in shared code still affects both. That is the price of the small step.

The ladder is worth knowing in order, because each rung costs roughly ten times the last: a bounded queue plus a separate thread pool inside one process, then a second deployment of the same artefact, then a separate service with its own data. Take the smallest rung that moves the incident count, and take the next one when evidence arrives rather than when someone predicts it.

Why the other options fail

  • Own service, own datastore, own team. This is the right answer when the two workloads have different data lifecycles, different release cadences or genuinely different owners. Here it pays for a data migration and a new network interface in order to solve a CPU-sharing problem, and it will not land before the next three incidents.
  • CPU-based autoscaling on the shared deployment. It scales the thing that is already saturated, and every new instance accepts both kinds of work, so checkout requests still queue behind image jobs. It also converts a visible incident into an invisible cost increase, which is how this choice survives for a year.
  • Raising the p99 target to 400 ms. This makes the breach disappear while leaving the customer experience exactly as it was, and it destroys the only signal that was telling you something is wrong. Moving a target is legitimate only when the old number was never justified by a business outcome, which is not the case for a checkout path.

What would flip the decision

If this changes Choose Because
Image work needs a GPU or a different base image Separate service The deployment unit must differ anyway
The two workloads need different release cadences Separate service One release stream is now the constraint
Image volume is small and bursts last seconds Bounded queue in one process A second deployment is not yet worth the baseline cost
Checkout must run in a region or compliance zone image work cannot Separate deployment per zone first The boundary is regulatory rather than architectural

When this is the wrong answer

If the saturated resource is the database rather than CPU, a second deployment of the same artefact makes it worse by adding connections. Find the saturated resource before choosing the rung: the fix is isolation of that resource, which may be a connection pool quota or a separate replica.