Embedding Pipeline Service  ·  View 02 of 22  ·  Context and scope

High-Level Architecture

The six stages between a document edit and a retrievable result, and the one place the cost of an edit is decided.

Editable source SVG draw.io All views
Capture Change capture feeds + sweep Change log Kafka, replayable Prepare Extract + normalise Tika, Tesseract Normalised text MinIO, 30 days Versioned chunker content hashing Decide what changed Chunk ledger PostgreSQL Ledger diff 75% reuse Embed Vector cache Redis Embedding fleet Ray Serve, GPU Index Index builder per contract Vector index Qdrant, immutable Alias etcd pointer Serve Retrieval API p99 180 ms Lexical index OpenSearch ACL filter fails closed cache miss alias fallback Embedding Pipeline Service — High-Level Architecture Interface / broker Queue / topic Application we own Data store Decision point Security / platform synchronous failure / alternate The chain is the build path. One contract change rebuilds from the ledger and the retained text, never from the source systems. v 1.0 · owner Data & AI Platform Architecture · date 2026-10

Decisions

  • The chunk ledger diff sits between preparation and embedding, so the GPU only ever sees chunks whose text actually changed. Assumed reuse on a typical edit is 75%.
  • The vector cache is keyed by (chunk hash, contract), which makes a reverted edit, a copied page and a shared template free rather than cheap.
  • Retrieval reads an alias, never an index name. That indirection is what makes a model migration an alias write.

Why this shape

  • Extraction, chunking and embedding are separated because their bottlenecks are CPU, memory and GPU and they do not saturate together.
  • Normalised text is retained for 30 days so a chunker change re-chunks without re-fetching 400 million documents from source systems.
  • The lexical index is not a nicety: it is the degradation path when the vector tier is unavailable.

Risks

  • If real-world chunk reuse is materially below 75%, the embedding fleet is undersized and the cost model is wrong. Reuse rate is therefore a first-class metric, not a curiosity.
  • The retained-text window is a bet that most re-chunking happens within a month of an edit. Past it, a rebuild costs a source-system load spike.