Walk me through the stress test you would run on a payments gateway in the mould of Razorpay before a festival sale, and tell me what the pass criterion is.
Show the full answer Hide the answer
What the interviewer is testing
Whether you know that a stress test produces a failure shape rather than a number, and whether you define recovery. Most candidates locate a breaking point and stop, which answers half the question the business asked.
The clarifying questions that change the answer
- Is the peak scheduled or discovered? A known sale window changes the test from "how much can we take" to "how fast do we get back".
- What degradation is acceptable to the business? Refusing new payment attempts is usually survivable; losing track of an authorised payment is not. That ordering decides where shedding is allowed to happen.
- Is the offered load open-loop? Merchant servers retrying on timeout make load rise as the gateway slows, which no closed-loop harness will reproduce.
- What is the recovery objective? How long may the gateway stay degraded after load returns to normal? If nobody has a number, that is the finding.
The arc of a strong answer
Ramp to 100% of forecast peak and hold until the signals are flat. Then step up in increments of 20% until a stated failure criterion trips — p99 beyond the deadline, error rate above budget, or any queue growing without bound. Record the breaking point and the first component to fail, which is rarely the one people predicted.
Then cut load back to 60% of normal and start a clock. The number that matters is overload recovery time: wall-clock from load removal until all of these hold at once — p99 inside the SLO, every queue at steady-state depth, no retry backlog, breakers closed, pools healthy, and no manual intervention used.
Why recovery is the criterion and not the breaking point
The arithmetic is the argument. A backlog drains at the spare rate, not the service rate. Suppose 20 minutes of overload accumulated a surplus of 8,000 messages per second against a service rate of 6,000, leaving roughly 2.4 million queued. If normal traffic leaves 2,000 messages per second of spare capacity, that backlog needs about 20 minutes to clear — and the sale lasts 90.
So a gateway that breaks at 140% of peak and recovers in 30 seconds is safe for the sale, and one that breaks at 220% and takes 20 minutes to settle is not. The breaking point is the number teams report and the recovery time is the number that decides whether the event survives.
The mechanisms that shorten recovery are structural, not operational: bounded queues, because you cannot drain a backlog you never accepted; deadline-aware dropping, so expired work is discarded instead of served; retry budgets at every caller; and shedding that disengages automatically with hysteresis so the system does not oscillate between shedding and overload.
Common weak answers
- "It should degrade gracefully." Unfalsifiable. Name the first thing to shed, the signal that triggers it, and the threshold.
- "Find the breaking point." Half the job. Two systems with identical breaking points can have recovery times an order of magnitude apart.
- "Autoscaling handles it." Autoscaling moves the breaking point and does nothing to the drain arithmetic, because the backlog still clears at the spare rate and new instances arrive minutes late.
What a strong answer adds
For a payments gateway the shedding point must sit at admission, before money moves. Shedding mid-authorisation converts an overload into reconciliation work that outlives the sale by days, which is a worse outcome than a rejected request. And the deliverable is one page: breaking point, first component to fail, recovery time, and the specific switch to flip — because during the sale nobody reads a report.
When this is the wrong answer
A stateless service with no internal queues and no uncontrolled client retries recovers the moment load stops. Measure it once, confirm it is seconds, and spend the effort elsewhere.