A team proposes replacing its weekly 48-hour soak with a 30-minute soak in every pipeline run. Review the proposal. What would you keep, what would you change, and what can a 30-minute run never find?
Show the full answer Hide the answer
What is actually required
The soak exists for one class of defect: anything that grows with elapsed time or cumulative operations rather than with load. Load tests vary the rate; soaks vary the clock. Replacing one with a shorter version of itself keeps the ceremony and loses the mechanism, so the question is which parts of the clock can be compressed.
The measurement problem, with numbers
A leak of 40 MB per hour produces 20 MB of growth in 30 minutes. Against a heap whose normal sawtooth is several hundred megabytes, 20 MB is inside the noise and no threshold will see it. This is why short soaks pass: not because the defect is absent but because the signal is below the resolution of the test.
The fix is to stop measuring the level and measure the slope. Fit a regression to post-collection live-set samples, and gate on a slope whose confidence interval excludes zero, projected to the release interval: "40 MB per hour over a 168-hour release cycle is 6.7 GB, which passes the 4 GB container limit in about 100 hours" is falsifiable. "Memory looked fine" is not.
What to change so the short run earns its place
Accelerate the ageing rather than the clock. The things that accumulate are created and destroyed by operations, so drive the operations:
- Churn, not just rate. Open and close connections, sessions, prepared statements, temporary files, cache entries and timers at 50–100× the production per-hour count. Steady request rate against a warm pool exercises none of this.
- Compress every period you control. Token lifetimes, TTLs, key rotation, log rotation, certificate renewal, leader re-election — set them to minutes in the test build so each occurs many times inside the window.
- Gate on a short list of monotone counters: post-collection live set, open file descriptors, thread or task count, connection count, prepared-statement count, registered timers, and disk used by the service. Every one should have a slope indistinguishable from zero.
Change the weekly soak's output too. Pass or fail is the wrong report; a slope table is the one people can act on, because it names the counter and the projected date it breaches a limit.
What a 30-minute run can never find
- Periods you cannot compress. A monthly partition roll, a 30-day token you cannot shorten without changing the code under test, a quarterly certificate.
- Volume-driven drift. A query plan that flips when a table passes some tens of millions of rows is a function of data, and no clock trick creates the data.
- Fragmentation. Allocator and index fragmentation need many allocation generations, and the pattern matters as much as the count.
- The deploy-free interval itself. This is the one teams miss: the pipeline restarts the process, so every counter starts at zero and the run always measures a freshly started service — the state production is almost never in.
What I would keep untouched
The weekly 48-hour run on the release candidate, for the uncompressible part, and the nightly restart if one exists. Removing the restart is a separate change with its own blast radius, and it should go only after the slope table shows zeros.
When not to run a soak at all
If the service is a stateless request handler whose processes are replaced by a deploy several times a week, the deploy cadence is the soak and no process ever lives long enough for slow drift to matter. There the honest answer is to delete the soak and put the same seven counters on a production dashboard with a slope alert, which costs nothing and watches the real workload.