Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
5 to work through
-
beginner Multiple choice
Users say the assistant's answers lack context, so a team proposes raising chunk size from 400 tokens to 1200 across a 60000-document policy corpus. Retrieval still returns the top 5 chunks. What is the dominant effect of that change?
3 min answer -
intermediate
A live assistant serves 1.8M documents chunked at a fixed 512 tokens with no overlap. You need to move to structure-aware chunks of about 900 tokens with 15% overlap and a newer embedding model, with no downtime and no quality regression. What is the sequence?
3 min answer -
intermediate
Users report the assistant gets numbers wrong when answering from documents containing tables. Diagnose.
2 min answer -
advanced
A retrieval system over code and documentation performs poorly. Which chunking decisions matter, and what is the common mistake?
2 min answer -
advanced
Review this ingestion pipeline. An HR assistant covers 4200 policy documents totalling about 9 million tokens. Ingestion runs a semantic chunker, generates three hypothetical questions per chunk and embeds those too, ensembles two embedding models, builds a knowledge graph of entity links, and adds a parent-document store. The nightly rebuild takes 11 hours and one engineer maintains all of it. Retrieval quality has never been measured. What would you remove and what would you keep?
3 min answer
3 terms in this topic
Chunk Boundary Strategy
How source documents are split for embedding, which determines whether retrieved passages are self-contained and coherent.
patternContextual Retrieval
Prepending a short generated description of where a chunk sits in its document before embedding it, so that a passage full of pronouns and bare figur…
practiceSemantic Chunking
Splitting documents along their meaning and structure rather than at fixed character counts, because retrieval quality is bounded by chunk quality.
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.