intermediate 3 min answer

A payments change passes its per-PR ephemeral environment and staging. Twenty minutes after the production deploy checkout p99 goes from 300 ms to 9 s and the database shows lock waits on one table. The ephemeral environment ran the same image and the same migration. What failed and which decision made the failure possible?

ephemeral environmentsparitylock contentiontest datapipeline
Show the full answer Hide the answer

The trigger

The change took a row lock earlier in the transaction and held it across a call that used to happen before the lock was taken. With one request in flight this is invisible: the lock is held for a few milliseconds and released. With forty concurrent workers hitting the same narrow key range, holding it across an external call makes the lock wait queue grow, each waiter holds a connection, the pool saturates, and latency for everything on that pool follows. The deploy did not break correctness; it changed the duration of a critical section.

Why the environment could not show it

Ephemeral per-PR environments are usually built for correctness, and they reproduce the axes that are easy to copy: same image, same schema, same configuration. They systematically miss the two axes where this class of defect lives — concurrency and data shape. The environment had one user and a few hundred seeded rows. There was no second transaction to contend with, no key skew to concentrate contention, and no query plan change from a table large enough for the optimiser to care.

That is a structural blind spot, not an oversight: a single-user environment with toy data cannot exhibit lock contention, connection-pool exhaustion, or a plan flip. Detection lagged in production for a related reason — the tests assert outcomes, and there was no assertion about how long anything held a lock.

The structural fix versus the tempting one

The tempting fix is a bigger, more production-like shared staging environment, which reintroduces the queue and the drift that ephemeral environments were built to remove, and still runs one user at a time.

The structural fix is two additions to the environment the team already has:

  • A production-shaped data subset. Not full volume: anonymised data that preserves row counts within an order of magnitude on the tables that matter and preserves skew in the top keys. The hot key is the point; a uniform synthetic dataset hides exactly this defect.
  • A 60-second concurrency smoke test at 10 to 20 requests per second against the ephemeral environment, with pass conditions on lock wait time, pool utilisation and p99 rather than on correctness. It costs a minute of pipeline time and it is the only thing in the pipeline that can fail for this reason.

The alert that would have caught it earlier

In production, lock wait time and connection-pool saturation per service, alerted on rate of change after a deploy rather than on an absolute threshold. p99 latency is the symptom everybody watches and the slowest to attribute; pool saturation names the service and the deploy.

Common weak answers

  • "Add an index." Sometimes the right fix, and here the plan was fine; the critical section was long.
  • "Load-test in staging before release." A weekly load test does not run on the change that introduced the regression, which is why the check belongs in the per-change environment.
  • "Roll back and investigate." Correct as an incident action and no answer to the question. The decision that made this possible was accepting an environment whose fidelity on concurrency and data shape was zero, and then treating a green pipeline as evidence about production behaviour.