1. RAG Architecture advanced

    A retrieval-augmented system serves millions of documents with sub-100 ms latency while embeddings are regenerated nightly. How should index rebuilds, hybrid retrieval, filtering and embedding versions be handled?

    3 min answer ragvector-searchindex-rebuildhybrid-retrieval
  2. RAG Architecture advanced

    A user asks the internal assistant a question and receives content from a document they cannot access. How did this happen and how is it prevented?

    2 min answer ragauthorizationsecurity
  3. RAG Architecture advanced

    An internal AI assistant gives confidently wrong answers. The team wants to upgrade to a better model. What do you check first?

    2 min answer airagretrievalevaluation
  4. Reranking intermediate

    A retrieval system's first-stage results are mediocre. Is reranking the right investment, and what does it cost?

    2 min answer rerankingretrievallatencycost
  5. Reranking advanced

    A support assistant adds a cross-encoder reranker. Offline NDCG at 10 rises from 0.61 to 0.74 and it ships. Two weeks later, answers drawn from long runbook pages are worse than they were before reranking, while short FAQ answers improved. Nothing errors and no alert fires. What happened?

    3 min answer rerankingcross-encodertruncationevaluation
  6. Reranking intermediate

    Retrieval returns 100 candidates and a cross-encoder reranks all of them. The service must hold 300 queries per second at a 400 ms p95 budget for the whole retrieval stage. Roughly what does that rerank cost in hardware and time and does it change the design?

    3 min answer rerankingcross-encodercapacity-planninggpu
  7. Tool Calling advanced

    A commerce platform exposes tools for a model to call. How should the tool interface be designed, and what differs from designing an API for developers?

    2 min answer tool-callinginterface-designerrorssafety
  8. Tool Calling beginner

    A team gives an assistant five tools. A task that needs four tool calls costs roughly six times a plain answer and takes about 9 seconds end to end, and the team expected the tools to be nearly free. What actually happens on each tool call?

    3 min answer tool-callingagent-looplatencytoken-cost
  9. Tool Calling advanced

    An agent can query the customer database, send emails and issue refunds. What is your security design?

    2 min answer agentssecurityauthorization
  10. Tool Calling intermediate

    An agent resolves support tickets in an average of 8 tool calls. A downstream inventory tool that used to answer in 80 ms starts taking 6 s but still returns correct results. Nothing errors and no circuit breaker trips. What happens, second by second?

    3 min answer tool-callingagentstimeoutsbackpressure
  11. Vector Databases advanced Multiple choice

    A multi-tenant assistant searches one HNSW index of 200M vectors and filters results to the requesting tenant. Large tenants are fine. For tenants with only a few thousand documents recall collapses and latency triples. Which change fixes it?

    3 min answer vector-databasehnswmulti-tenancyfiltering
  12. Vector Databases advanced Multiple choice

    A platform needs similarity search over a large and growing corpus. What decides between a dedicated vector database, a vector extension to an existing store, and a search engine with vector support?

    2 min answer vector-databaseoperational-costfilteringscale