1. Chunking & Retrieval intermediate

    A live assistant serves 1.8M documents chunked at a fixed 512 tokens with no overlap. You need to move to structure-aware chunks of about 900 tokens with 15% overlap and a newer embedding model, with no downtime and no quality regression. What is the sequence?

    3 min answer chunkingre-embeddingmigrationshadow-traffic
  2. Chunking & Retrieval advanced

    A retrieval system over code and documentation performs poorly. Which chunking decisions matter, and what is the common mistake?

    2 min answer chunkingretrievalstructurecontext
  3. Chunking & Retrieval advanced

    Review this ingestion pipeline. An HR assistant covers 4200 policy documents totalling about 9 million tokens. Ingestion runs a semantic chunker, generates three hypothetical questions per chunk and embeds those too, ensembles two embedding models, builds a knowledge graph of entity links, and adds a parent-document store. The nightly rebuild takes 11 hours and one engineer maintains all of it. Retrieval quality has never been measured. What would you remove and what would you keep?

    3 min answer ragingestionover-engineeringevaluation
  4. Chunking & Retrieval intermediate

    Users report the assistant gets numbers wrong when answering from documents containing tables. Diagnose.

    2 min answer ragchunkingquality
  5. Chunking & Retrieval beginner Multiple choice

    Users say the assistant's answers lack context, so a team proposes raising chunk size from 400 tokens to 1200 across a 60000-document policy corpus. Retrieval still returns the top 5 chunks. What is the dominant effect of that change?

    3 min answer chunkingretrievalembeddingsprecision
  6. Embeddings advanced

    A marketplace uses embeddings for product similarity. The embedding model is upgraded. What is the operational consequence, and how should the system be designed for it?

    2 min answer embeddingsmigrationversioningreindex
  7. Embeddings beginner Multiple choice

    A support search drops any result scoring below 0.75 cosine similarity. The team re-embeds the whole corpus with a newer embedding model and rebuilds the index; now the same cutoff throws away almost every result. What is the mistake?

    3 min answer embeddingssimilaritycalibrationthresholds
  8. Embeddings advanced

    You are building retrieval over an organisation's documents. Which three decisions most affect quality?

    2 min answer embeddingschunkinghybrid-retrievalevaluation
  9. Guardrails advanced

    A platform must prevent harmful outputs in a user-facing AI feature. Where should guardrails sit, and what does each layer catch?

    2 min answer guardrailslayersmoderationfalse-positives
  10. Guardrails intermediate

    An AI feature needs guardrails on inputs and outputs. Where should they run, and what is the latency and reliability consequence?

    2 min answer mistralguardrailssafetylatency
  11. Guardrails advanced

    An insurance company wants an LLM to draft claim decision letters. What is your architecture?

    2 min answer ai-governanceguardrailshuman-in-the-loop
  12. Guardrails intermediate Multiple choice

    Your output safety classifier blocks 0.3% of responses in English and 6% in Portuguese and Turkish. Complaints arrive only from those two markets. The vendor confirms the same model version serves every region and reports no incident. Where do you look first?

    3 min answer guardrailsclassifierscalibrationlocalisation