1. Performance & Capacity advanced

    A distributed training job runs at 90% scaling efficiency on 512 GPUs. At 1,024 GPUs it drops to 60%. Walk me through where the time went, what you would measure, and what you would try.

    3 min answer nvidiadistributed trainingcollectivesnccl
  2. Performance & Capacity intermediate

    A search engine holds its index entirely in memory to guarantee predictable latency. What does that buy, what does it cost, and when is it the wrong choice?

    2 min answer typesensemeilisearchin-memorypredictability
  3. Performance & Capacity advanced

    Millions of customers attempt to buy a small number of items at a scheduled instant. Which performance constraint dominates, and why do conventional scaling techniques fail against it?

    2 min answer flash-salecontentionhot-keyinventory
  4. Performance & Capacity intermediate

    Peak trading day is six weeks away and expected to be four times normal traffic. What do you do in those six weeks?

    3 min answer capacityload-testingpeakreadiness
  5. Performance & Capacity intermediate

    Precompute every user's timeline at write time, or assemble it at read time? Explain why the answer for a social feed is neither.

    2 min answer case-studytwitterfan-outpower-law
  6. Performance & Capacity intermediate

    Your system handles 1,000 requests per second today. Marketing says a campaign will bring 10,000 next month. What breaks first, and how do you find out?

    2 min answer capacitybottleneckload-testingscaling
  7. Profiling & Optimisation intermediate Multiple choice

    A colleague spent a week optimising a function and the endpoint is 2% faster. What went wrong in the approach?

    2 min answer optimisationamdahlprofilingmethod
  8. Profiling & Optimisation intermediate

    A team has profiled a service and found the top three functions by CPU time. Why is optimising them often the wrong next step?

    2 min answer profilingoptimisationamdahlmeasurement
  9. Profiling & Optimisation advanced

    Discord's Read States service, written in Go, showed latency spikes every two minutes like clockwork. The team had written it carefully with very few allocations, and the spikes appeared regardless of load. Their published account from 2020 explains the cause and the rewrite that followed. What was happening, and what does it teach about periodic latency?

    3 min answer discordgarbage collectiontail latencyruntime
  10. Queueing Theory intermediate

    A capacity review shows services running at 85% CPU. Finance suggests raising it to 95% to save money. What is your response?

    2 min answer capacitylatencycost
  11. Queueing Theory advanced Multiple choice

    A platform runs its worker fleet at 85% average utilisation to control cost. Queue times are becoming unpredictable. What does queueing theory say is happening?

    2 min answer browserstackqueueingutilisationlatency
  12. Queueing Theory advanced

    An inference service's utilisation rises from 70% to 90% and latency more than triples. Why is the relationship non-linear, and what does that imply for capacity planning?

    2 min answer queueing-theoryutilisationlatencyvariability