1. Capacity Modelling advanced

    A training cluster has 512 NVIDIA H100-class GPUs in 64 eight-GPU nodes and the scheduler reports 96 GPUs free. Roughly how much collective bandwidth per GPU does a new 64-GPU job get, and what should the capacity report say instead of "96 GPUs free"?

    3 min answer nvidianvlinkncclfragmentation
  2. Capacity Modelling intermediate Multiple choice

    Alibaba reported a peak of 583,000 order creations per second during Singles' Day in 2020. Suppose one promoted item accounts for 0.5% of those orders and its stock is held in a single row. Roughly how many updates per second does that row have to absorb?

    2 min answer alibabaflash salehot keywrite contention
  3. Capacity Modelling intermediate Multiple choice

    You run in three availability zones and must survive losing one. What is your maximum normal utilisation and why?

    2 min answer capacityheadroomfailure-domainsutilisation
  4. Capacity Modelling intermediate Multiple choice

    Zoom's CEO wrote on 1 April 2020 that maximum daily meeting participants went from roughly 10 million at the end of December 2019 to more than 200 million in March 2020. For any platform absorbing a 20× demand rise over three months rather than three minutes, which capacity decision does the slower surge force that the fast one does not?

    3 min answer zoomcapacity-modellinglead-timequota
  5. Concurrency advanced Multiple choice

    A platform must handle millions of concurrent connections per process cluster where most connections are idle most of the time. Which concurrency model fits, and what does the wrong choice cost?

    2 min answer concurrencyevent-loopthreadsconnections
  6. Concurrency beginner Multiple choice

    A service handles 100 requests per second with a pool of 10 worker threads. The team raises the pool to 100 threads. Throughput stays at roughly 100 rps and p99 latency gets much worse. What is the primary reason?

    2 min answer concurrencylittles lawbottleneckthread pools
  7. Concurrency advanced Multiple choice

    A service holds a concurrency limit of 60 with a 50 ms mean service time and serves clients that give up after 2 seconds. Offered load triples. Where should the excess requests wait?

    3 min answer concurrencybounded-queuedeadlinesgoodput
  8. Concurrency advanced Multiple choice

    Why is limiting concurrency more effective than limiting request rate when protecting a service?

    2 min answer concurrencyrate-limitinglittles-lawprotection
  9. Connection Pooling intermediate

    A service scales its application instances from 20 to 200 and the database begins refusing connections. What happened, and what is the correct architecture?

    2 min answer connection-poolingproxyscalingdatabase-limits
  10. Connection Pooling advanced

    A service under load has high latency and low CPU. The team increases the database connection pool and it gets worse. Why?

    2 min answer sharechatconnection-poolcontentionqueueing
  11. Connection Pooling intermediate

    A service works fine at 10 instances. At 60 instances during a peak, the database starts refusing connections. Explain the arithmetic.

    2 min answer poolsscalingdatabases
  12. Connection Pooling beginner

    Why does putting a connection pool in front of a database make a service faster when the database does exactly the same query work either way?

    3 min answer connection-poolingtls-handshakepostgresconcurrency