1. Queueing Theory advanced

    WeChat's DAGOR overload control, published at SoCC 2018, profiles a server's load from the average waiting time of requests in its pending queue rather than from CPU utilisation, sheds by business and user priority, and propagates admission levels between services. Why is queuing time the right signal, and why does the propagation matter more than the shedding?

    3 min answer tencentwechatdagoroverload control
  2. Soak Testing intermediate

    A platform passes every load test yet degrades after several days of continuous operation. What class of problem is this, and what testing finds it?

    2 min answer soak-testingmemory-leaksresource-exhaustionendurance
  3. Soak Testing intermediate

    A service is restarted nightly by a cron job and nobody remembers why. What do you suspect and how do you confirm it?

    2 min answer soak-testingmemory-leakresource-leaktrends
  4. Soak Testing intermediate

    A service performs well in load tests and degrades after several days in production. What class of problem is this, and how is it found before release?

    2 min answer credsoak-testingleaksfragmentation
  5. Soak Testing intermediate

    A team proposes replacing its weekly 48-hour soak with a 30-minute soak in every pipeline run. Review the proposal. What would you keep, what would you change, and what can a 30-minute run never find?

    3 min answer soak-testingdriftregression-gatememory
  6. Stress Testing advanced

    At 150% of capacity, what should your service do? Describe the mechanisms.

    2 min answer overloadload-sheddingadmission-controlprioritisation
  7. Stress Testing advanced

    Walk me through the stress test you would run on a payments gateway in the mould of Razorpay before a festival sale, and tell me what the pass criterion is.

    3 min answer stress-testingoverloadrecoveryqueue-drain
  8. Stress Testing intermediate

    What is the purpose of stress testing beyond load testing, and what should a communication platform learn from deliberately breaking itself?

    2 min answer stress-testingbreaking-pointdegradationrecovery
  9. Tail Latency advanced

    A fashion marketplace's average latency is excellent but p99 is unacceptable for a small, commercially important customer segment. How should tracing, segmentation, queue analysis and dependency timing guide the investigation?

    2 min answer myntratail-latencyp99segmentation
  10. Tail Latency advanced

    A request fans out to 100 services in parallel. Each responds within 10 ms for 99% of calls. What fraction of user requests are slow, and what do you do?

    2 min answer fan-outlatencyhedgingarithmetic
  11. Tail Latency advanced

    A video platform's request path fans out to twenty backend services in parallel. Each has a p99 of 50 ms. Why is the overall p99 far worse than 50 ms, and what fixes it?

    2 min answer tail-latencyfan-outhedgingamplification
  12. Tail Latency advanced

    You add hedged requests to cut tail latency. It works well, then during a traffic peak the service collapses. Explain.

    2 min answer hedgingoverloadlatency