pattern

Bulkhead

Partitioning resources so that exhaustion in one area cannot consume the capacity another area depends on.

isolationresilienceresources

Named after ship compartments, and the analogy is exact: the point is not to prevent flooding but to confine it.

The failure it prevents is the most common shape of cascading outage. A service uses one thread pool or connection pool for all its outbound calls. One dependency slows down, its calls occupy the pool, and the service can no longer serve requests that never touch that dependency. A partial failure has become a total one.

Bulkheading assigns separate pools per dependency, so a slow dependency exhausts only its own allocation. The same principle appears at every scale: separate connection pools per downstream, separate thread pools per workload class, separate compute for interactive and batch, separate clusters per tenant tier, and cell-based architecture at the top of the range where entire independent stacks each serve a slice of users.

The cost is utilisation — partitioned resources idle while another partition is saturated, so the total capacity needed is higher than a shared pool would require. That is the trade, stated plainly: you are buying containment with efficiency.

The sizing question people get wrong: pools must be sized so that all pools together fit within the process's real limits. Ten pools of 100 connections against a database that accepts 400 has created a different exhaustion problem one layer down.