RAG Architecture
Retrieval, grounding, citation and the permissions RAG can enforce.
6 to work through
-
intermediate
A business unit wants an assistant answering questions from 200,000 internal documents. They ask whether to fine-tune a model or use retrieval. How do you decide?
2 min answer -
advanced
A RAG system returns fluent answers that are frequently wrong. Where in the pipeline is the problem most likely, and how would you find out?
2 min answer -
advanced
A retrieval-augmented system over enterprise documents gives confidently wrong answers. Which failure modes are responsible, and where should engineering effort go?
2 min answer -
advanced
A retrieval-augmented system serves millions of documents with sub-100 ms latency while embeddings are regenerated nightly. How should index rebuilds, hybrid retrieval, filtering and embedding versions be handled?
3 min answer -
advanced
A user asks the internal assistant a question and receives content from a document they cannot access. How did this happen and how is it prevented?
2 min answer -
advanced
An internal AI assistant gives confidently wrong answers. The team wants to upgrade to a better model. What do you check first?
2 min answer
4 terms in this topic
Hybrid Retrieval
Running lexical keyword search and dense vector search together and fusing the results, because each fails where the other succeeds.
conceptRetrieval Grounding
Constraining a model's answer to content retrieved from an authoritative corpus, with citations and an abstention path - so that output quality becom…
patternRetrieval Pipeline Stages
The stages that turn a user question into grounded context — query processing, retrieval, reranking and assembly — each independently tunable.
practiceRetrieval-Generation Separation
Measuring whether the correct passage was retrieved, separately from whether the answer was correct - the single diagnostic that turns unfalsifiable …
Neighbouring topics
AI-Era Architecture
General material on architecting systems that include models.
LLM Application Architecture
The shape of a production system with a model in the request path.
Vector Databases
Approximate nearest-neighbour search, filtering and re-indexing.
Embeddings
Dense representations, model coupling and the migration they imply.
Chunking & Retrieval
Structure-aware splitting, hybrid search and why chunking dominates quality.
Reranking
Cross-encoders improving precision more than a bigger embedding model.
Model Selection
Capability, latency, cost and the evaluation that decides between them.
AI Gateways
Centralised routing, keys, quotas, caching, logging and safety policy.
Prompt & Version Management
Prompts as reviewed, versioned, evaluated production configuration.
Agent Architectures
Loops, planning, memory and the boundaries an agent must not cross.
Tool Calling
Typed tool interfaces, narrow parameters and per-tool authorisation.
Multi-Agent Systems
Coordination, hand-off and whether more agents actually help.
LLM Evaluation
Held-out sets, rubric judging, CI gates and production sampling.
AI Observability
Logging prompts, versions, retrieved context and cost per request.
Guardrails
Deterministic checks on input and output that fail closed.
Prompt Injection Defence
Breaking the private-data, untrusted-input, outbound-channel combination.
AI Cost Management
Token accounting, routing, caching and the context-window budget.
Human in the Loop
Gating by reversibility and blast radius, and avoiding approval fatigue.
ML Platform
Feature stores, training pipelines, registries and deployment.