1. Peak Event Readiness advanced

    A trading platform faces its largest load at a precisely known time every day. What does readiness for a scheduled peak involve beyond capacity?

    2 min answer zerodhapeakreadinesswarmup
  2. Peak Event Readiness advanced

    Retail leadership proposes a six-week change freeze covering the entire estate from mid-November. Respond.

    2 min answer retailfreezeriskgovernance
  3. Peak Event Readiness advanced

    You expect 10x traffic on a known date six weeks away. Walk through the preparation.

    3 min answer peakcapacitychange-freezeload-shedding
  4. Performance & Capacity advanced

    A distributed training job runs at 90% scaling efficiency on 512 GPUs. At 1,024 GPUs it drops to 60%. Walk me through where the time went, what you would measure, and what you would try.

    3 min answer nvidiadistributed trainingcollectivesnccl
  5. Performance & Capacity advanced

    Millions of customers attempt to buy a small number of items at a scheduled instant. Which performance constraint dominates, and why do conventional scaling techniques fail against it?

    2 min answer flash-salecontentionhot-keyinventory
  6. Profiling & Optimisation advanced

    Discord's Read States service, written in Go, showed latency spikes every two minutes like clockwork. The team had written it carefully with very few allocations, and the spikes appeared regardless of load. Their published account from 2020 explains the cause and the rewrite that followed. What was happening, and what does it teach about periodic latency?

    3 min answer discordgarbage collectiontail latencyruntime
  7. Queueing Theory advanced Multiple choice

    A platform runs its worker fleet at 85% average utilisation to control cost. Queue times are becoming unpredictable. What does queueing theory say is happening?

    2 min answer browserstackqueueingutilisationlatency
  8. Queueing Theory advanced

    An inference service's utilisation rises from 70% to 90% and latency more than triples. Why is the relationship non-linear, and what does that imply for capacity planning?

    2 min answer queueing-theoryutilisationlatencyvariability
  9. Queueing Theory advanced

    WeChat's DAGOR overload control, published at SoCC 2018, profiles a server's load from the average waiting time of requests in its pending queue rather than from CPU utilisation, sheds by business and user priority, and propagates admission levels between services. Why is queuing time the right signal, and why does the propagation matter more than the shedding?

    3 min answer tencentwechatdagoroverload control
  10. Stress Testing advanced

    At 150% of capacity, what should your service do? Describe the mechanisms.

    2 min answer overloadload-sheddingadmission-controlprioritisation
  11. Stress Testing advanced

    Walk me through the stress test you would run on a payments gateway in the mould of Razorpay before a festival sale, and tell me what the pass criterion is.

    3 min answer stress-testingoverloadrecoveryqueue-drain
  12. Tail Latency advanced

    A fashion marketplace's average latency is excellent but p99 is unacceptable for a small, commercially important customer segment. How should tracing, segmentation, queue analysis and dependency timing guide the investigation?

    2 min answer myntratail-latencyp99segmentation