Review this retrieval setup. An assistant over a policy corpus retrieves the top 5 chunks. Users complain it cites the 2023 edition of a policy. Inspection shows the corpus holds about seven near-identical revisions of every document and the top 5 are usually five revisions of the same page. Which change would you make first?
Show the full answer Hide the answer
What is actually required
The assistant must answer from the policy in force, with a citation a compliance reader can follow. Nothing in the requirement asks for history on the default path.
What this setup is actually doing
Seven revisions of one page are seven vectors within noise distance of each other, so one document can occupy every slot in the top 5. The effective breadth of retrieval has collapsed from five sources to one, and the context window is spending five times the tokens to deliver one passage. The version that wins a slot is arbitrary — whichever revision happens to phrase a sentence marginally closer to the query — which is exactly why it sometimes lands on 2023.
Two independent defects are stacked here: a corpus-identity defect (nothing says which revision is current) and a retrieval-diversity defect (nothing prevents one source filling the slate). Fixing the second without the first leaves the system able to cite a superseded policy.
The one change that matters
Give every chunk the key (document_id, revision, chunk_index), keep only the current revision in the index the assistant queries, and move superseded revisions to a separate archive index reachable only when a query asks about history. Retrieval quality improves twice over: the citation is correct by construction, and five slots again hold five different documents.
Why the other options fail
- Raise top-k from 5 to 20. The intuitive move, and the most expensive wrong one. Twenty slots fill with roughly the same duplicate families, so you pay four times the context tokens, push time-to-first-token up, and add the well-documented effect of diluting a long context with near-identical passages. The wrong revision is still in the running.
- Add a cross-encoder reranker. A reranker reorders what it is given; it cannot add what was never retrieved, and it has no idea which revision supersedes which. It will faithfully rank seven copies of the same text. It is a reasonable second investment once the candidates are distinct.
- A larger embedding model. Near-identical texts stay near-identical in any embedding space. This re-embeds the whole corpus, invalidates the calibrated score thresholds, and leaves the defect untouched.
What I would leave alone
The 512-token chunks and the top-5 budget. Both are fine, and both look like the problem because the symptom shows up in them. Changing the knob nearest the symptom is how a duplicate-identity bug becomes a six-week chunking project.
When this is the wrong answer
For a legal, audit or claims-history corpus, the superseded revision is often the answer — what did the policy say on the date of the incident. There, deduping destroys the product. The fix becomes a required effective-date filter in every query, a prompt that always states which revision it used, and dedupe applied only within one revision rather than across them.