1. Bottleneck Analysis advanced

    A marketplace's search page is slow. CPU is moderate, memory is fine, and the database reports healthy query times. How do you find the bottleneck?

    2 min answer bottleneckuslqueueingmeasurement
  2. Bottleneck Analysis advanced

    A service's p99 latency has tripled over three months with no single obvious change. How do you investigate?

    2 min answer performanceprofilingmethodregression
  3. Bottleneck Analysis advanced

    CPU, memory, disk and network all look healthy, and the service will not go faster. What are you looking for?

    2 min answer bottlenecklogical-resourcessaturationcontention
  4. Bottleneck Analysis advanced

    p99 latency on the booking endpoint jumped from 300 ms to 4 s at 09:00 today. Error rate is normal. Walk through your diagnosis.

    2 min answer debugginglatencydiagnosismethod
  5. Caching for Performance advanced

    A search platform adds a cache and sees only a small latency improvement. What are the likely reasons, and what should be measured?

    2 min answer cachinghit-rateworking-setzipf
  6. Caching for Performance advanced

    What happens to your system if the entire cache tier is flushed at peak traffic?

    2 min answer cachecold-startcapacityfailure-modes
  7. Caching for Performance advanced

    Which caching failure appears only under load, and what are the fixes in order of deployability?

    2 min answer cache-stampedestale-while-revalidatecoalescingjitter
  8. Capacity Modelling advanced

    A compute platform must plan accelerator capacity where lead times are months, demand is uncertain, and jobs vary from minutes to weeks. How should the capacity model be built?

    2 min answer capacity-modellinglead-timeuncertaintyscheduling
  9. Capacity Modelling advanced

    A game hosts a scheduled in-world event for 30 million concurrent players. How should matchmaking, instance allocation, pre-scaling, login queues and degradation be planned, and what is rehearsed?

    3 min answer epic-gamesfortnitelive-eventsmatchmaking
  10. Capacity Modelling advanced

    A training cluster has 512 NVIDIA H100-class GPUs in 64 eight-GPU nodes and the scheduler reports 96 GPUs free. Roughly how much collective bandwidth per GPU does a new 64-GPU job get, and what should the capacity report say instead of "96 GPUs free"?

    3 min answer nvidianvlinkncclfragmentation
  11. Concurrency advanced Multiple choice

    A platform must handle millions of concurrent connections per process cluster where most connections are idle most of the time. Which concurrency model fits, and what does the wrong choice cost?

    2 min answer concurrencyevent-loopthreadsconnections
  12. Concurrency advanced Multiple choice

    A service holds a concurrency limit of 60 with a 50 ms mean service time and serves clients that give up after 2 seconds. Offered load triples. Where should the excess requests wait?

    3 min answer concurrencybounded-queuedeadlinesgoodput