Embedding Pipeline Service · View 02 of 22 · Context and scope
Decisions
- The chunk ledger diff sits between preparation and embedding, so the GPU only ever sees chunks whose text actually changed. Assumed reuse on a typical edit is 75%.
- The vector cache is keyed by (chunk hash, contract), which makes a reverted edit, a copied page and a shared template free rather than cheap.
- Retrieval reads an alias, never an index name. That indirection is what makes a model migration an alias write.
Why this shape
- Extraction, chunking and embedding are separated because their bottlenecks are CPU, memory and GPU and they do not saturate together.
- Normalised text is retained for 30 days so a chunker change re-chunks without re-fetching 400 million documents from source systems.
- The lexical index is not a nicety: it is the degradation path when the vector tier is unavailable.
Risks
- If real-world chunk reuse is materially below 75%, the embedding fleet is undersized and the cost model is wrong. Reuse rate is therefore a first-class metric, not a curiosity.
- The retained-text window is a bet that most re-chunking happens within a month of an edit. Past it, a rebuild costs a source-system load spike.