Embedding Pipeline Service — Architecture One-Pager
Solution Architecture v1.0 · open source, self-hosted on Kubernetes · Data & AI Platform Architecture · 2026-10 · 22 views · 17 architecture decision records
Embedding Pipeline Service · Solution Architecture v1.0 · open source, self-hosted on Kubernetes · Data & AI Platform Architecture · 2026-10 · 22 views · 17 architecture decision records
An index is identified by its embedding contract and is immutable under it. Changing the contract builds a new index beside the old one and cuts over by alias; it never updates vectors in place.
Every workspace product people use daily has grown the same feature: a search box that understands what you meant, and an assistant that answers from your own documents with citations. The visible surface is a box. What decides whether it works is a pipeline nobody sees, keeping embeddings in step with a corpus that is edited eight million times a day, across twenty-five thousand organisations, under permissions the platform does not own. The hard part is not building it once. It is that embedding models improve every few months while a corpus lives for years — so a design in which a model change is a project will be two years out of date within two years, and the retrieval quality the entire product rests on will decay with no moment at which anyone decided to let it.
A shared internal platform. Product teams register a corpus and declare a freshness tier; the platform captures changes, extracts and chunks the text, embeds only what changed, builds immutable per-contract indexes behind mutable aliases, and serves hybrid retrieval with serve-time permission filtering to fourteen product surfaces. Open source throughout, self-hosted on Kubernetes: Kafka as the change log, PostgreSQL as the chunk ledger, MinIO for retained text and snapshots, Redis as the vector cache, a KubeRay GPU fleet for inference, Qdrant and OpenSearch for retrieval, etcd for aliases, Argo Workflows for rebuilds, ClickHouse for quality and drift.
What it is, and what it is not
- An index identified by its embedding contract, immutable under it — not An index as a long-lived mutable store that is upgraded in place, document by document
- Chunk identity derived from content, so an edit costs what it changed — not Chunk identity derived from position, so a paragraph insert re-embeds the document
- The corpus as the truth, with every vector reproducible from retained text plus a pinned contract — not The index as the truth, with vectors as precious state that has to be preserved
- Permission resolved at serve time against the authority, failing closed — not Permission copied into the index at build time and refreshed when someone remembers
- Revocation on the read path in seconds, indexing in minutes — not Revocation inheriting the pipeline's latency because it travels the same path
- Cutover gated on measured retrieval quality — not Cutover gated on the build being finished
The decisions that are the architecture
- The embedding contract is the identity of an index, not configuration applied to one (ADR-01) — A contract is (normaliser version, chunker version, model digest, pooling rule, dimensionality). A vector's primary key is (chunk hash, contract id), and the index builder refuses a vector whose contract does not match the index's. "We upgraded the model" therefore cannot mean a gradual, invisible, half-migrated result — the most common way these systems rot.
- Build beside, then flip (ADR-02) — A new contract builds a shadow index that receives the live change stream alongside the serving one. Cutover is an alias write; rollback is the same write in reverse, against an index that is still built, still fed and still measured. The storage doubling and the dual-write window are budgeted capacity states with a declared 21-day life, not incidents.
- No query ever spans two contracts (ADR-03) — Similarity scores from different models are not comparable and cannot honestly be normalised. Admitting that early is what makes progressive per-tenant cutover a real option and query-time fan-out across two models a trap, and it is why the alias is per tenant rather than global.
- The corpus is the truth; the chunk ledger is ours (ADR-04) — Content and permissions stay with their owner. The platform's own system of record is the chunk ledger, and every index is a projection rebuildable from the ledger, the retained text and a pinned contract. That is what makes a corrupt index, a bad build and a wrong chunker one remedy instead of three disaster procedures.
- Chunk identity is a content hash, so an edit costs what it changed (ADR-05) — A typical 41-chunk document edit produces three changed chunks. Everything downstream — GPU time, index writes, cost — is proportional to that number rather than to the document. Chunk reuse rate is therefore tracked as an operational signal, because it is the single number that sets the cost of the edit stream.
- The chunker version is a contract element (ADR-06) — A chunker improvement is as much a full rebuild as a model change, and pretending otherwise produces an index whose chunk boundaries vary by when a document was last touched. Making the chunker a contract element prices chunking experiments honestly and keeps every stored vector reproducible.
- Three freshness lanes, with bulk on preemptible capacity (ADR-09) — Interactive, standard and bulk, each with its own published batch wait. GPU inference is far cheaper per vector in large batches, and a large batch is bought with waiting — so freshness is spent deliberately, lane by lane, and a 14-day rebuild is affordable because bulk is preempted rather than interactive being queued behind it.
- A document version becomes visible atomically (ADR-10) — The new version is retrievable only once every chunk of it is indexed; until then the previous version serves. A half-reindexed document matches on old and new wording at once and no reader can tell which they are seeing, which is strictly worse than being stale.
- Permission is resolved at serve time and fails closed (ADR-13) — The platform asks the corpus's authority whether this subject may read this document, on every retrieval, and withholds everything when the authority does not answer. It is the one hard hot-path dependency in the design, accepted because the alternative — an ACL copy inside the index — is a second source of truth for the most consequential fact in the system.
- Revocation travels the read path, not the pipeline (ADR-14) — A deletion or newly-restricted document stops being retrievable within five seconds through a suppression list consulted before the permission batch, while full erasure propagates through every store behind it. Revocation must not inherit the indexing latency it has nothing to do with.
- Model replicas are addressed by digest and gated by a frozen probe (ADR-15) — A registry serving different weights behind the same name would silently invalidate every stored vector and present as "search got worse" months later. Every rollout embeds a frozen reference set and compares against stored expected vectors; divergence fails the rollout, not the pipeline.
- Cutover is gated on measured quality, and drift is monitored in three places (ADR-16) — Recall@10 against a frozen evaluation set must not fall more than two percentage points against the outgoing contract. Corpus drift, model change and query drift all present as "search got worse" and only one is fixed by rolling back, so each is monitored separately and every alert says which the evidence supports.
- Snapshots of a rebuildable store exist for recovery time, not durability (ADR-17) — A from-scratch rebuild takes three to fourteen days and satisfies no recovery-time objective anyone would sign, so indexes are snapshotted to reach a 4-hour RTO — and a restored snapshot is not trusted until the deletion log has been replayed over it, because otherwise a restore resurrects an erased document.
Why this should still hold in ten years
The decisions above are deliberately stated in terms of authority, identity and reproducibility rather than in terms of products. That is what should let them survive the thing most likely to change about this system, which is every component in it.
- The boundary is not a technology. "An index is identified by its embedding contract and is immutable under it" says nothing about Qdrant, Kafka, PostgreSQL or Ray. A different vector store, a different log, a managed embedding API or an architecture not yet invented can each satisfy it, and everything downstream — dual-write, alias cutover, per-tenant progressive flip, rebuild as routine — follows from the boundary rather than from the stack.
- It is designed for the one change that is certain. Embedding models will keep improving on a cadence of months. A design whose central operation is "adopt a new model" gets better as the field does; a design whose central operation is "serve the model we chose" gets worse. The cost the architecture pays — reproducibility, retained text, a rebuild budget — is the premium on that option, and it is paid once rather than per migration.
- Reproducibility outlives any particular index. Because every vector is derivable from retained text plus a pinned contract, the platform can answer "what exactly produced this result" years later, and can move to a store that does not exist yet by rebuilding rather than migrating. The chunk ledger is the asset; the indexes are inventory.
- The permission decision ages well. Resolving access at serve time against the owner's authority costs latency now and will keep costing it. But sharing models grow more complex over time, never less, and a platform that copied an ACL model in 2026 would be maintaining a divergent replica of a model nobody documented by 2030. Asking is slower and stays correct.
- The failure posture does not depend on scale. Degrade a dimension rather than the service; declare staleness rather than imply freshness; fail closed on access and fail soft on everything else. These hold at a tenth of this scale and at ten times it, which matters because a platform's traffic is the assumption most likely to be wrong.
- What will date. The specific numbers — 768 dimensions, 4,000 chunks a second, a 30-day text window, a 2-point recall gate — are all products of 2026 hardware and 2026 models and should be expected to move by an order of magnitude. They are stated as assumptions precisely so that moving them is an edit rather than an excavation. Quantisation, learned chunking and multi-vector retrieval are the three places where a change in the field would most plausibly force a structural revision rather than a numeric one.
Non-functional targets
Every target below is a stated assumption. They are included because a requirement without a number cannot be designed against, and because each of these drove at least one decision in the record.
| Quality | Target | How it is met | View |
|---|---|---|---|
| Retrieval availability | ≥ 99.95% monthly on the retrieval path | Stateless retrieval API across three zones; an index already built serves with the entire build path stopped | 07 |
| Retrieval latency | p50 ≤ 45 ms, p99 ≤ 180 ms in-region for a hybrid top-50 over a tenant partition; p99 ≤ 400 ms above 50 M chunks | Tenant-partitioned ANN index, query embedding on the same contract, batched ACL resolution | 15 |
| Freshness | Interactive lane p50 ≤ 30 s, p95 ≤ 2 min, p99 ≤ 10 min; standard p95 ≤ 30 min; bulk none | Lane admission at capture, per-lane batch wait, oldest-unembedded age as both SLI and scaling trigger | 14 |
| Revocation | ≤ 5 s at p99 for a deleted or newly-restricted document to stop being retrievable | Read-path suppression list checked before the permission batch, independent of indexing latency | 10 |
| Ingest availability | ≥ 99.9% monthly, measured as change events durably logged and acknowledged | Accept returns on durability in the change log; never waits on extraction or inference | 13 |
| Control-plane availability | ≥ 99.5% monthly | Deliberately lower: an outage freezes cutovers and widens freshness, and removes no retrieval | 09 |
| Throughput, live stream | 3,000 chunks/s sustained, 5× burst for 15 minutes | Autoscaled Ray Serve replicas on on-demand GPU capacity, lane-isolated | 14 |
| Throughput, rebuild | 4,000 chunks/s background (≈ 14 days) or 18,500 chunks/s surge (≈ 72 h at ≈ 4.5× hourly cost) | Bulk lane on preemptible GPU capacity, spilling to on-demand; fleet sized on this, not on 280/s mean | 14 |
| Durability, authoritative stores | RPO 0, RTO 15 min for the chunk ledger, contract registry, index catalogue and deletion log | PostgreSQL with Patroni, synchronous standby in a second zone; append-only deletion log | 11 |
| Recovery, indexes | RTO 4 h from snapshot; a from-scratch rebuild is 3 to 14 days and satisfies no RTO | Periodic index snapshots to MinIO, with the deletion log replayed over any restore before it serves | 11 |
| Quality gate | Recall@10 must not fall more than 2 percentage points against the outgoing contract; quantisation ≤ 1 point | Frozen per-corpus evaluation set scored by the quality harness; the gate blocks the alias flip | 19 |
| Cost | ≤ $0.45 per million chunks embedded at background rate; ≤ $11 per million chunks per month to store and serve; full re-embed ≤ $2,200 GPU | Chunk reuse rate and cache hit rate monitored as operational signals; migrations priced before they start | 18 |
Scope
In scope
- Corpus registration, change capture from feeds and webhooks, and a reconciliation sweep that finds what the feed missed
- Extraction, OCR and versioned normalisation of the product's core formats, with offsets preserved back to the source document
- Versioned chunking with content-hash chunk identity, and a chunk ledger that diffs document versions
- Embedding inference on an operated GPU fleet addressed by model digest, with a vector cache and a rollout probe
- Immutable per-contract vector indexes behind mutable aliases, with hybrid lexical retrieval and atomic version visibility
- Serve-time permission filtering that fails closed, and a read-path suppression list giving 5-second revocation
- Embedding contracts, dual-write migration, alias cutover and rollback, and migrations priced before they start
- A frozen evaluation set per corpus, drift monitoring in three places, and per-tenant cost showback
Explicitly out of scope
- The generative answer layer: prompting, synthesis, citation rendering and the conversation state behind them
- The document stores, file stores and connected SaaS corpora themselves, which remain systems of record
- The permission model: the platform resolves access against the owner's authority and never decides it
- The product user interfaces, including any indexing-in-progress signal a surface may want to show
- Choosing the embedding model, which arrives as a digest and becomes an input to a contract
- Multimodal corpora, learned chunking, multi-region index serving and lazy long-tail re-embedding, all deferred to Phase 3
What to build first, and what it has to prove
The MVP is not a smaller version of the platform. It is the smallest system that can demonstrate the critical design decision is both true and affordable, because everything else in the package is downstream of that.
- One corpus, one tenant, change-feed ingestion with monotonic versions over a durable change log
- Extraction with offset mapping and retained normalised text; a versioned chunker with content-hash identity
- A chunk ledger that diffs document versions, with the reuse rate instrumented from day one
- An embedding fleet addressed by model digest, a vector cache, and the frozen reference probe on rollout
- One ANN index per tenant partition per contract, behind an alias, with atomic version visibility
- Serve-time permission filtering that fails closed, plus the read-path suppression list
- A second contract, built beside the first, dual-written, gated on recall@10, cut over and rolled back once on purpose
- Edit a paragraph in a 300-chunk document. Prove that the number of chunks embedded is in single figures, and that the document is retrievable under its new text inside the interactive budget.
- Revert that edit. Prove that no inference happened at all, because every chunk hash was already in the cache.
- Register a second contract with a different model digest, build it beside the first, and cut over. Prove the flip is a single write and that retrieval never served results from both contracts at once.
- Roll the cutover back. Prove it is the same single write, and that the previous index was still current because dual-write never stopped.
- Revoke access to a document that is in the serving index. Prove it stops being retrievable inside five seconds, before any pipeline work has run.
- Stop the entire build path — capture, extraction, inference — and prove retrieval still meets its p99 against the index already built.
- Delete an index and rebuild it from the ledger and retained text alone, with no source system involved, and measure how long it took. That number is the migration budget.
Open risks, carried rather than hidden
| Risk | If it lands | Response |
|---|---|---|
| Real-world chunk reuse on edit is materially below the assumed 75% | The embedding fleet is undersized, the cost model is wrong by the same factor, and the interactive freshness budget is missed under normal load rather than under burst | Instrument reuse rate per corpus from the first day of the MVP and treat it as a launch criterion, not a report. If a corpus's editing pattern rewrites whole documents, that corpus needs a different chunking strategy or a different freshness tier — and the number tells you which before a product has shipped on it. |
| The frozen evaluation set is unrepresentative, so the cutover gate passes a contract that is worse in production | A migration completes, the alias flips, and retrieval quality falls for queries nobody wrote into the set. The rollback window is 21 days and the regression may not be noticed inside it | Build the set from the real query log rather than from imagination, refresh a sampled slice of it each quarter, and pair the gate with implicit signals — answer acceptance, reformulation rate — watched for two weeks after each wave. Treat a gate pass with a falling acceptance rate as a failed migration. |
| The permission authority becomes the platform's availability ceiling | Retrieval is contractually 99.95% but is in practice bounded by a system the platform does not own and cannot fail open on. Every authority incident is a retrieval incident | Keep the ACL check batched and cacheable for the life of a request, publish the authority's own availability beside the retrieval SLO so the dependency is honest, and negotiate a target with its owners. Revisit pre-filtering only with measured recall loss in hand — an ACL copy trades an availability problem for a correctness one. |
| A learned or semantic chunker becomes clearly better, and chunking stops being deterministic | The reproducibility principle weakens: a vector can no longer be derived from retained text plus a pinned contract unless the chunker model is itself pinned and retained like an embedding model | Treat a learned chunker as a model from the outset — pinned by digest, probed on rollout, a contract element with the same rebuild cost. If that is unaffordable for a given corpus, the honest answer is that the corpus keeps a deterministic chunker, not that reproducibility is relaxed. |
| The dormant long tail never migrates, and the estate permanently spans contracts | Storage for superseded indexes is never reclaimed, two or three contracts are maintained indefinitely, and the cost model that justified the architecture quietly inverts | Make lazy re-embedding an explicit Phase 3 decision with a retirement deadline attached, not a drift. Measure the fraction of each corpus ever queried, and if the tail is large, dematerialise idle indexes rather than leaving them on an old contract — an index nobody queries should cost storage, not a maintained contract. |
The reasoning behind every component and technology choice is in the Architecture Decision Record: 17 records across 6 areas, each with the alternatives that lost and what the choice costs.