Fastly's June 2021 incident saw a valid customer configuration trigger a latent bug from a May deployment, returning errors across most of its network for under an hour. You run a multi-tenant API where one tenant's input can reach a latent defect in shared code. Which isolation approach limits the damage?
Show the full answer Hide the answer
The deciding property
The failure in the stem is a poison pill: one tenant's valid input deterministically triggers a defect in shared code. That word "deterministically" settles the choice, because it means the failure follows the input rather than occurring at random.
So the question becomes: what bounds the set of tenants whose requests can reach the same copy of the defective code? Only an approach that partitions tenants into independent units with their own deployments of the full stack does that.
Why cells
A cell is a complete, independently-deployed copy of the stack serving a fixed subset of tenants. Tenant 400's poison configuration reaches cell 3's instances, and cells 1, 2 and 4 through 10 never execute that code path with that input. With ten cells, the blast radius is bounded at roughly 10% of tenants, by construction, for any input-triggered defect.
Cells give a second property that matters more than the isolation on the day: they are the unit of deployment. The May deployment that introduced the latent bug would have reached one cell first, so the bake period in a single cell is what converts a latent defect into a contained one. Cells bound the blast radius of both bad input and bad code, which is why they are the structural answer rather than a mitigation.
What it costs
Substantial, and worth stating: N copies of the stack with N× the fixed overhead, a tenant-to-cell routing layer that must itself be highly available and is now a shared component, cross-cell operations (analytics, admin, a tenant that must move) becoming genuinely hard, and N deployments to track so release engineering gets more complex rather than less. Cells are a serious architectural commitment and are usually introduced after an incident makes the case.
Why the other options fail
- Multiple regions. The near-miss. Regions isolate infrastructure failure — a power event, a network partition, a regional service outage. They do not isolate a logic defect, because every region runs the same code, and the poison configuration is distributed to all of them within the config propagation window. This is the most common wrong answer and the reason multi-region deployments still have global outages.
- Shuffle sharding. A genuinely good technique and the wrong one for a deterministic defect. Assigning each tenant a random subset of, say, 2 of 8 workers makes it unlikely that two tenants share the same full set, which bounds the damage from a noisy tenant — one consuming disproportionate resources. But a poison pill crashes or corrupts every worker it reaches, and because the tenant's requests spread across its assigned subset, the defect propagates to that whole subset and overlapping tenants are affected too. Shuffle sharding dilutes resource contention; it does not contain a deterministic crash.
- A circuit breaker around the config processor. Treats the symptom after the fact. A breaker opens once failures are observed, which is after the input has already triggered the defect wherever it reached, and it cannot tell a poison input from a dependency problem. Useful as defence in depth, not as isolation.
When cells are the wrong answer, and what would flip the decision
For a team of eight with 200 tenants, cells are the wrong recommendation even though they are the structurally strongest option. N copies of the stack, a highly-available routing tier and N deployments to track is a multi-quarter commitment that consumes the reliability budget of a small platform entirely. Below roughly a few thousand tenants, or wherever the operational headcount cannot carry N deployments, take staged config rollout plus fast rollback and revisit when an incident makes the case.
| If this changes | Choose | Because |
|---|---|---|
| The risk is a noisy tenant exhausting shared resources | Shuffle sharding | Cheap, no stack duplication, and it is exactly what that technique bounds |
| The risk is infrastructure failure, not logic | Multi-region or multi-AZ | Independent failure domains for the correlated hardware and network cases |
| Fewer than a few dozen tenants, or one tenant is most of the revenue | Dedicated single-tenant stacks | Cells converge on this anyway, with less machinery |
| Early stage, small team | Staged config rollout plus fast rollback | Catches most of this for days of work instead of quarters |
What a strong answer adds
That the cheapest mitigation for the specific incident described is not an isolation architecture at all: staged rollout of configuration, treating config as a deployment artefact. Fastly's own account emphasises rapid rollback, and a config change that reaches 1% of the fleet for ten minutes before going wider contains most of this class for a fraction of the cost of cells. Knowing which fix is proportionate to the stage of the company is the senior judgement being tested, and recommending cells to a team of eight is the wrong answer even though cells are the structurally strongest one.