Review this ingestion pipeline. An HR assistant covers 4200 policy documents totalling about 9 million tokens. Ingestion runs a semantic chunker, generates three hypothetical questions per chunk and embeds those too, ensembles two embedding models, builds a knowledge graph of entity links, and adds a parent-document store. The nightly rebuild takes 11 hours and one engineer maintains all of it. Retrieval quality has never been measured. What would you remove and what would you keep?
Show the full answer Hide the answer
What is actually required
Nine million tokens at 400-token chunks is about 22000 chunks. That is a small corpus. A single embedding model over 22000 vectors fits comfortably in memory on one machine, rebuilds in minutes, and answers in single-digit milliseconds. Every component in this pipeline exists to improve ranking within those 22000 items, and not one of them has been shown to improve anything, because nothing is measured.
The multipliers are the problem. Three hypothetical questions per chunk plus the chunk itself is 4 vectors per chunk. Two embedding models doubles it again. The index holds roughly 176000 vectors for a corpus that needs 22000 — an 8x amplification bought on faith, and the 11-hour rebuild is the bill.
What I would do first, before removing anything
Write 80 labelled questions from real HR tickets and measure recall at 10. One afternoon of work. That number is the only thing that can justify or condemn any component here, and it also sets the ceiling: if the correct passage reaches the candidate set 95 percent of the time, no amount of downstream cleverness is the constraint, and the effort belongs in generation or prompting instead.
What I would remove
- The two-model ensemble. Scores from different embedding models are not on a comparable scale, so fusing them requires rank-based combination, which the team almost certainly has not implemented correctly. It doubles index size and rebuild time for a gain that, on a homogeneous single-language corpus, is typically small. Remove it and keep whichever model measures better alone.
- The knowledge graph. Entity graphs earn their cost on multi-hop questions: "which policies changed after the 2024 restructure and who approved them". HR questions are overwhelmingly single-hop lookups. Until the labelled set contains multi-hop questions that the graph answers and vector search does not, this is an unused index with a maintenance burden.
- Hypothetical question generation. It improves matching when user phrasing differs sharply from document phrasing, and it triples the index. A short generated context prefix on each chunk gets most of the same benefit at a quarter of the storage, and keeps one vector per chunk so scores stay interpretable.
What I would keep, even though it looks odd
The parent-document store stays. It is cheap, it is not in the vector index at all, and it solves the real problem the other components were bought to solve: match precisely on a small child chunk, then hand the model the whole section so the answer has its surroundings. Keep the semantic chunker too, but pin its version, because a chunker that silently changes its splits between releases makes every quality comparison meaningless.
The change that matters more than any removal
An 11-hour rebuild means a policy correction published at 09:00 is not answerable until the next day. For an HR assistant, that is the actual defect, and it is caused by the amplification rather than by the corpus. After the cuts, incremental ingestion of changed documents takes minutes, which turns a nightly batch into a near-real-time pipeline and removes the single-engineer bus factor at the same time.
When this is the right pipeline
At 50 million documents with mixed languages, multi-hop compliance questions and a measured recall ceiling below 0.85, several of these components start paying. The distinguishing evidence is always the same: a labelled set showing where the pipeline currently fails. Sophistication added before measurement is not architecture, it is hope with a rebuild window.