1. Little's Law advanced

    A downstream call goes from 50 ms to 3 seconds at constant traffic. Use Little's Law to explain what happens to the caller and why services that never call it also fail.

    2 min answer littles-lawcascading-failurebulkheadspools
  2. Little's Law intermediate

    A platform must size worker pools for a peak of 50,000 requests per second with a 200ms average service time. How does Little's Law inform the answer, and what does it not tell you?

    2 min answer dream11littles-lawconcurrencycapacity
  3. Little's Law advanced

    A service's latency rises sharply while throughput plateaus and CPU sits at 50%. Use queueing theory to explain what is happening and what to measure next.

    3 min answer queueing-theorylittles-lawconcurrencyutilisation
  4. Little's Law intermediate Multiple choice

    An API platform has an average of 400 requests in flight and processes 2,000 requests per second. What is the average latency, and how is Little's Law used for capacity decisions?

    2 min answer littles-lawconcurrencycapacityconnection-pools
  5. LLM Application Architecture advanced

    A workspace product adds AI features over user content. Which architectural decisions dominate, and which are commonly deferred at cost?

    2 min answer llm-applicationpermissionslatencycaching
  6. LLM Application Architecture advanced

    An LLM chat product serves millions of multi-turn conversations. How do KV-cache reuse, prefix caching, continuous batching and session affinity change the architecture, and what breaks when a session lands on a different GPU?

    3 min answer character-aikv-cacheprefix-cachinginference
  7. LLM Application Architecture beginner Multiple choice

    Your chat endpoint streams tokens to the browser. A request fails after 300 of an expected 500 tokens are already on the user's screen. The team wants to apply the same automatic retry policy the rest of their API uses. Why is this different?

    3 min answer streamingretriesidempotencylatency
  8. LLM Application Architecture advanced

    Zoom publicly describes the architecture behind its AI Companion as a federated approach - its own models used alongside third-party frontier models, with work routed by task rather than every request going to a single provider. What problem does that structure solve that a single-provider design does not, and where would copying it be a mistake?

    3 min answer model-routingmulti-providercostevaluation
  9. LLM Evaluation advanced

    A platform ships AI features and cannot tell whether changes improve or degrade quality. What evaluation infrastructure is required, and in what order?

    2 min answer evaluationregressionofflineonline
  10. LLM Evaluation advanced

    A team ships an LLM feature and cannot tell whether changes improve it. What evaluation infrastructure is needed, and what does it not solve?

    2 min answer unacademyevaluationgolden-setregression
  11. LLM Evaluation advanced

    You are asked to prove an AI assistant is good enough to launch. How do you construct the evidence?

    2 min answer evaluationlaunchgovernance
  12. LLM Evaluation intermediate

    You want an LLM-as-judge suite to gate every pull request that touches a prompt. It holds 400 cases; each case runs the product (about 3000 input and 400 output tokens) then a judge call (about 3500 input and 200 output). Roughly what does one run cost and how long does it take, and does the answer change the gate design?

    3 min answer llm-evaluationci-gatecostrate-limits