Contextual Retrieval
also called Contextual Chunk Prefixing, Contextual Embeddings
Prepending a short generated description of where a chunk sits in its document before embedding it, so that a passage full of pronouns and bare figures still matches the query that should find it.
A chunk reads "The company's revenue grew by 3% over the previous quarter." Nobody can tell which company, which quarter or which filing. The embedding cannot either, so a query naming the company and the quarter ranks that chunk somewhere in the middle of the corpus, and the assistant answers from a worse passage.
This is the dominant retrieval failure in enterprise corpora, and it is not caused by a weak embedding model. It is caused by splitting: the information that made the sentence interpretable lived in the document title, the section heading and the preceding paragraph, and chunking threw all three away. Contextual retrieval puts a small amount of it back, in the chunk itself, before the vector is computed.
Why it matters
Retrieval quality is the ceiling on a RAG system. No prompt and no larger model recovers a passage that never reached the context. Contextual prefixing attacks the failure at the stage where the information was lost, which is why it outperforms downstream fixes that only reorder what was already found.
It also survives the usual objection to ingest-time cleverness: it changes one vector per chunk rather than multiplying the index, so scores stay interpretable and the rebuild stays cheap.
Implementation patterns
- Generate 50 to 100 tokens of context per chunk naming the document, the section and whatever the chunk's pronouns and figures refer to, and prepend it to the chunk text before embedding.
- Use a cheap model with the whole document cached as the prompt prefix. Anthropic reports a one-time cost of $1.02 per million document tokens under prompt caching, which makes the technique affordable for corpora in the tens of millions of tokens.
- Index the prefixed text in both the vector index and the lexical index. The published gain comes from doing both: contextual embeddings alone cut the top-20 retrieval failure rate from 5.7% to 3.7%, and adding contextual BM25 took it to 2.9%.
- Store the original chunk separately and pass that, not the prefixed version, to the generator - the prefix is a retrieval aid, not content the model should quote.
- Version the prefix generation. A changed prompt or model produces different prefixes and therefore different vectors, so the prefix version belongs in the index metadata alongside the embedding model version.
Industry example
Anthropic's 2024 engineering write-up on contextual retrieval gives the numbers above from its own evaluation: a baseline top-20 failure rate of 5.7%, 3.7% with contextual embeddings, 2.9% with contextual BM25 added, and 1.9% when reranking is layered on top - a 67% reduction in failed retrievals overall. The same post is explicit that for knowledge bases under about 200000 tokens, roughly 500 pages, the whole retrieval apparatus can be skipped by putting the corpus in the prompt.
Failure scenarios
- The generated context hallucinates a date or an entity, and the chunk is now retrievable by a query about something it does not contain.
- The prefix dominates the embedding. If the context is long relative to a short chunk, every chunk from one document collapses towards the same vector and within-document ranking is destroyed.
- Ingest cost is underestimated because prompt caching was not configured, turning a one-dollar-per-million-token job into a fifty-dollar one.
- The prefix leaks into answers, because the retrieved text passed to the generator was the prefixed version, producing citations of sentences that do not exist in the source.
- Prefixes become stale when the document is edited and only the changed chunk is re-generated, so neighbouring chunks describe a structure that has moved.
Trade-offs
| Choose contextual retrieval | Gains | Pays |
|---|---|---|
| Corpus of self-referential documents | Large measured recall gain at one vector per chunk | An LLM pass over the whole corpus at ingest |
| Frequently queried, rarely changed corpus | Cost is one-time and amortised | Re-generation on every document edit |
| Mixed lexical and semantic queries | Helps both indexes at once | Two indexes to keep in step |
The honest comparison is against reranking, which needs no ingest change and helps immediately. They are complementary rather than alternatives, and the published figures combine them.
When not to use it
When the corpus fits in the context window, put it in the prompt and delete the pipeline. When chunks are already self-describing - API reference pages, product records, support tickets with structured headers - the prefix restates what is already there and buys nothing. When nobody has measured recall, this is a guess: build 60 to 100 labelled questions first, because the technique is worth roughly what your current boundary loss costs, and that varies enormously by corpus. And when documents change hourly, the re-generation cost recurs at a rate the one-time framing does not cover.
Interview question
Q: You have measured recall at 20 on your labelled set at 0.94 and answer correctness at 0.71. A colleague proposes contextual retrieval. What do you tell them?
What a strong answer covers: the recall ceiling is already 0.94, so at most 6 points of the failure is retrieval; the other 23 points sit in reranking, context assembly or generation, and contextual prefixing cannot touch them. The right move is to diagnose where the correct passage is lost after retrieval - is it in the candidate set but below the cut, or in the context and ignored - and spend there. Contextual retrieval becomes the priority when recall itself is the constraint.
Quick check
Quiz: Contextual retrieval reduced Anthropic's reported top-20 retrieval failure rate from 5.7% to what, with contextual BM25 included, and to what again with reranking added? 2.9%, then 1.9%.
Flashcard: Why does prepending generated context to a chunk beat buying a better embedding model? — Because the information was destroyed at chunking time, so no encoder can recover it from the chunk; the prefix puts it back before the vector is computed, at one vector per chunk.