At 09:07 UTC a CDN route recovers and a queued asset is released. Within seconds the API gateway sees roughly 1.5 million requests per second, about three times its normal peak. Gateway tasks start failing health checks and being replaced, though CPU never saturates and the upstream services look healthy. Canva published this incident for 12 November 2024. Where is the time going, and what do you look at first?
Show the full answer Hide the answer
The first three things I would look at, and why in that order
- Thread states, not CPU. CPU at 60% with requests timing out means threads are blocked, not busy. A thread dump — or a continuous profiler's wall-clock view, which is the one that shows blocked time — separates "working" from "waiting on a lock" in a way that no CPU graph can.
- The health-check path. Tasks are being killed. If the health endpoint shares a thread pool with request handling, a saturated pool fails its checks although the process is alive and progressing, turning a load problem into a capacity-destroying one.
- What the instrumentation itself is doing. This is the step people skip, and in this incident it is the answer.
The diagnosis
Canva's published post-incident review attributes the gateway's degradation under the surge to a telemetry bug that caused thread locking. The shape of the failure is worth stating generally, because it recurs:
An APM or metrics agent sits on the request path and takes a lock held across work proportional to load. A metric registry that synchronises on a map when a new tag combination appears, a span exporter whose queue is guarded by one mutex, a sampler that locks to update a counter. At normal rates the lock is held for microseconds. At three times peak, arrival rate crosses the point where the lock's service time exceeds its inter-arrival time, and the queue in front of it goes from empty to unbounded with no intermediate state — a queueing discontinuity, not a slope, which is why the dashboards show a cliff.
The failure then feeds itself. Blocked threads cannot answer health checks, the orchestrator replaces the tasks, the replacements start cold and immediately receive the same surge, and effective capacity falls while demand is still at 3x. The replacement loop makes recovery slower than doing nothing.
The misleading signal
CPU utilisation. It will sit at a comfortable number throughout, and it will be used in the incident channel as evidence that the gateway is fine and the problem is upstream. It is evidence of exactly the opposite: a saturated service whose CPU is low is a service that is blocked, and the gap between "requests in flight" and "CPU busy" is the measurement that should have been on the dashboard instead.
The second is the CDN. The routing problem was real and over by the time the gateway failed. Trigger and failure are separated in time, so a team anchored on "the CDN broke us" waits for a recovery that does not help.
The fix, in order
- Get the instrumentation off the critical section. Pre-register metric tags at startup rather than lazily on first use; make the span export queue lock-free or per-thread; make the agent's failure mode "drop telemetry" rather than "block the request". Telemetry must be the first thing shed under load, never the last.
- Separate the health-check path. A dedicated thread or executor for liveness, so a saturated request pool reports "unhealthy, do not route to me" without being killed. Distinguish readiness from liveness: shed traffic from a struggling task, do not destroy it.
- Admission control at the gateway. A concurrency limit above measured capacity returns fast errors to a fraction of users instead of slow failures to all of them, and keeps the fleet alive to recover.
- Break the herd at the source. Hundreds of thousands of clients unblocked on the same asset at the same instant. Retry jitter and a staggered release of a long-blocked resource turn a spike into a ramp.
The alert that would have caught it earlier
Not request rate, which is a symptom of the world rather than of the system. Alert on the ratio of in-flight requests to CPU-seconds consumed — a rising count of concurrent requests with flat CPU is the signature of lock contention, and it rises before the first timeout. Second: alert on health-check latency as a distinct series from request latency. When liveness checks start taking hundreds of milliseconds, the task-replacement spiral is seconds away and has not started yet.
When this is the wrong diagnosis
Blocked threads with low CPU have two likelier causes, and reaching for the agent first wastes the incident if either is the real one. A saturated downstream connection pool gives the identical signature. So does a synchronous call to a dependency whose latency has risen, where threads block on a socket read rather than a lock. The discriminator is in the thread dump: waiting on a monitor held by another application thread points inside the process; waiting on a socket or a pool permit points outward. Check the pool and the downstream first — likelier and cheaper to confirm — and suspect the agent when the blocked frames sit inside the telemetry library itself.
What a strong answer adds
That the observability stack is production code with production failure modes, and it is on the hot path of every request. An agent added to improve reliability reduced it, and no amount of dashboarding finds that, because the dashboards are drawn by the thing that is broken. The test is a load test run with instrumentation enabled at production fidelity — most load tests run with sampling turned down, which is precisely the configuration that cannot reproduce this.