A payments service adds a connection-pooling ambassador beside every pod. After rollout the database's connection count falls from 1800 to 120 and its CPU drops 30%, but p99 time-to-commit rises from 90 ms to 210 ms while p50 is unchanged at 12 ms. The ambassador logs no errors and sits at 8% CPU. Where is the time going, and what do you change first?
Show the full answer Hide the answer
The first three things I would look at
- The ambassador's own wait time — how long a client waits for a server connection — not its CPU or error rate. Waiting consumes no CPU and raises no error, so neither existing signal can see it.
- Pooling mode and pool size per pod, against peak concurrent transactions per pod. "Fewer connections" was the goal, and somebody had to pick a number.
- The distribution of transaction duration, specifically whether a small class of long transactions shares the pool with the short ones.
The diagnosis
The queue did not disappear; it moved out of the database and into the proxies, where nothing was watching it. Before, 1,800 connections contended inside the database. Now a few thousand application threads contend for 120 server connections, and the wait is paid at the start of the transaction, before any statement reaches the server.
The latency shape is the tell. p50 flat with p99 doubled means most requests find a free connection and a minority queue — the signature of a utilisation close to saturation, not of a slow dependency. At a mean transaction time of 8 ms, 120 server connections can in principle sustain around 15,000 transactions a second, so capacity looks fine on average; what sets the tail is per-pool utilisation. As a pool approaches full occupancy, waiting time grows roughly as utilisation over one minus utilisation, so at 90% occupancy the wait is about nine times the service time. One class of long holders is enough to push it there: a 400 ms reporting query occupying a connection in a four-connection pool blocks every short transaction behind it.
The misleading signal
Every database metric improved, which is exactly what the ambassador was bought for. Connection count, server CPU, memory per connection: all better, all irrelevant to the symptom, because the database cannot see clients queueing outside it. The proxy's 8% CPU misleads in the same way.
The fix, in order
- Export pool wait time and pool saturation per pod, and alert on them. Wait-time p99 becomes a first-class indicator; saturation above 80% for five minutes pages. Nothing else is diagnosable until this exists.
- Separate pools by statement class. A small dedicated pool for long transactions and reports, so they cannot occupy the pool that serves checkout.
- Size the pool from concurrency, not from thrift. Peak in-flight transactions per pod times a headroom factor of about two. With 40 concurrent transactions per pod, a four-connection pool is the defect.
- Only then consider transaction-level pooling for a further reduction, and audit what breaks first: session-scoped state — prepared statements, advisory locks, temporary tables, session variables — stops working when a connection is handed to a different client between transactions.
The alert that would have caught it earlier
Wait time at the pool, published by the proxy, with a p99 budget set as a fraction of the end-to-end latency target. A pooling layer without a wait-time metric is an unobservable queue in the hot path.
When this is the wrong answer
If p50 had risen along with p99, the pool is simply too small and resizing is the whole answer. If the proxy's CPU were saturated, you are looking at per-connection overhead instead and need more proxy, not a different pool shape. The flat-p50-with-raised-p99 shape is what points at queueing with occasional long holders.