# Embedding Pipeline Service

**Solution Architecture v1.0 · open source, self-hosted on Kubernetes · Data & AI Platform Architecture · 2026-10 · 22 views · 17 architecture decision records**

Every workspace product people use daily has grown the same feature: a search box that understands what you meant, and an assistant that answers from your own documents with citations. Notion's workspace search, Slack search, Dropbox Dash, Confluence AI search, Glean across all of them — the visible surface is a box, and the thing that decides whether it works is invisible. This package is that invisible thing, as a shared internal platform inside one collaboration-SaaS company: an assumed 25,000 customer organisations and 3 million monthly active users, 400 million documents, 4.8 billion live chunks, 8 million document versions edited a day, 12,000 retrieval queries a second — built on Kafka as the change log, PostgreSQL as the chunk ledger, MinIO for retained text and snapshots, Redis as the vector cache, a KubeRay GPU fleet addressed by model digest, Qdrant and OpenSearch for retrieval, etcd for aliases, Argo Workflows for rebuilds, and ClickHouse for quality and drift.

The product is not the ability to turn text into vectors. It is **the ability to change the model** — because embedding models improve every few months while a corpus lives for years, and a design in which a model change is a project will be two years behind within two years, with the retrieval quality the whole product rests on decaying at no moment anyone decided on.

The design rests on one rule: **an index is identified by its embedding contract — normaliser, chunker, model digest, pooling rule, dimensionality — and is immutable under it. Changing any element of the contract builds a new index beside the old one and cuts over by alias; it never updates vectors in place.**

The decisions that carry the design:

- **The contract is identity, not configuration.** A vector's primary key is (chunk hash, contract id), and the index builder refuses a vector whose contract does not match the index's. "We upgraded the model" therefore cannot mean a gradual, invisible, half-migrated index whose cosine similarities are being compared across two spaces — the most common way these systems rot (ADR-01).
- **Build beside, then flip.** A new contract builds a shadow index that receives the live change stream alongside the serving one, so "complete" means complete. Cutover is an alias write; rollback is the same write in reverse against an index still built, still fed and still measured. The storage doubling is a budgeted capacity state with a 21-day life, not an incident (ADR-02).
- **No query ever spans two contracts.** Similarity scores from different models are not comparable and cannot honestly be normalised. Admitting that is what makes progressive per-tenant cutover a real option and query-time fan-out across two models a trap (ADR-03).
- **The corpus is the truth; the chunk ledger is ours.** Content and permissions stay with their owners; the platform's own system of record is the chunk ledger, and every index is a projection rebuildable from the ledger, the retained text and a pinned contract. A corrupt shard, a bad build, a wrong chunker and a failed migration then have one remedy instead of four disaster procedures (ADR-04).
- **Chunk identity is content-derived, so an edit costs what it changed.** A typical 41-chunk document edit produces three changed chunks, and everything downstream is proportional to that number rather than to the document. Chunk reuse rate is tracked as an operational signal because it is the single number that sets the cost of the edit stream (ADR-05).
- **The chunker version is a contract element.** A chunking improvement is as much a full rebuild as a model change, and pretending otherwise produces an index whose chunk boundaries vary by when each document was last touched (ADR-06).
- **A reverted edit is free.** The vector cache is keyed by (chunk hash, contract) and scoped per tenant — scoped deliberately, because a cache shared across tenants would be an existence side channel on exact document text (ADR-07).
- **Normalised text is retained 30 days, so a rebuild never touches the source systems.** Rebuildability that depends on somebody else's API is a promise you do not control, and OCR paid once per version rather than once per rebuild is the dominant saving on a scanned corpus (ADR-08).
- **Three freshness lanes, with bulk on preemptible capacity.** GPU inference is far cheaper in large batches and a large batch is bought with waiting, so the wait is assigned per lane and published. That is what makes a 30-second interactive promise and a 14-day, $2,200 rebuild simultaneously achievable (ADR-09).
- **A document version becomes visible atomically.** Until every chunk of a version is indexed, the previous version serves — because a half-reindexed document matches on both old and new wording at once and no reader can tell which they are seeing, which is strictly worse than being stale (ADR-10).
- **At-least-once delivery, idempotence from our own keys.** The change log is partitioned by document so per-document order is guaranteed, and a monotonic source version arbitrates, so an out-of-order redelivery cannot overwrite newer text (ADR-11).
- **Verify what you were told, and never infer destruction from silence.** A nightly sweep finds the edits the feed missed; a document missing from a source inventory is a candidate requiring explicit confirmation, never a deletion — otherwise a paginating source bug erases a tenant's retrievable corpus (ADR-12).
- **Permission is resolved at serve time and fails closed.** The platform asks the corpus owner's authority on every retrieval and withholds everything when it cannot. It is the one hard hot-path dependency in the design, accepted because an ACL copied into the index is a divergent replica of the most consequential fact in the system (ADR-13).
- **Revocation travels the read path, not the pipeline.** Five seconds at p99 through a suppression list, while full erasure propagates behind it with per-store evidence — and any restored snapshot has the deletion log replayed over it before it serves, which closes the erasure hole most recovery procedures leave open (ADR-14, ADR-17).
- **Models are addressed by digest and probed on every rollout.** A registry serving different weights behind the same name would silently invalidate every stored vector and surface weeks later as "search feels worse". A frozen reference set turns that into a failed deployment (ADR-15).
- **Cutover is gated on measured quality, and drift is three problems with one symptom.** Recall@10 must not fall more than two points against the outgoing contract; corpus drift, model change and query drift are monitored separately, because only one of the three is fixed by rolling back (ADR-16).

The architecture one-pager (including why the design should still hold up in ten years, and the five risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~15 min) and [docs/decision-record.md](docs/decision-record.md) (~72 min).

---

## What is here

| Path | Contents |
|---|---|
| `ask.md` | The high-level requirements this architecture answers — capability sections, quantified NFRs, failure classes, eight open questions and the critical design decision |
| `diagrams/index.html` | The landing page: 22 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 17 decision records, the stack table and the glossary as markdown |
| `specs/part-*.json` | The authored view specs. Edit these, never the assembled `views.json` |
| `specs/manifest-*.json` | The authored acts and reasoning cards. Edit these, never the assembled `manifest.json` |
| `specs/adr-onepager.json`, `specs/adr-records-*.json` | The authored one-pager and decision records. Edit these, never the assembled `adr.json` |
| `scripts/build.sh` | Rebuilds every deliverable from the specs. Node 20+ and nothing else |
| `scripts/render-adr.mjs` | Renders the one-pager and decision record into the landing page and into `docs/` |

## Rebuilding

```bash
cd architecture/deliverables/use-cases/embedding-pipeline-service
bash scripts/build.sh
```

No draw.io Desktop, no browser, no network. The build runs generate → validate (strict) → route check → editable SVG → plain SVG → draw.io copy → HTML pages and index → ADR injection → link check, and stops at the first failure.

Current state: **22 views, 0 validator errors, 0 validator warnings, 182 edges routed with 0 routing errors and 8 clutter warnings, 443 icons embedded, 0 unresolved icons, all relative links resolve.** The eight clutter warnings are long `rel` labels fanning into the centre of the system-context and integration views, and the two gutter edges that carry the deletion path — each was reduced as far as it could be without losing a label the view needs.

## The seven acts

| Act | Views |
|---|---|
| Context and scope | 01 system context · 02 high-level architecture |
| People and journeys | 03 actors and journeys · 04 find what I just wrote · 05 ship a new corpus · 06 upgrade the embedding model |
| Structure | 07 layered architecture · 08 platform components · 09 integration surface |
| Data | 10 data flow · 11 storage zones · 12 data model |
| Runtime | 13 edit to retrievable · 14 embedding pipeline · 15 retrieval paths · 16 contract migration and cutover |
| Operations | 17 deployment architecture · 18 observability · 19 contract lifecycle |
| Assurance | 20 security trust zones · 21 retrieval authorisation · 22 failure classes |

## What is assumed

Every quantity in this package is a stated assumption, sized for a collaboration-SaaS company of the shape described above. None is measured from a production system. They are stated explicitly so they can be argued with: an explicit number forces the architecture to commit, and vagueness cannot be reviewed. The figures that set design boundaries are the 14-day background and 72-hour surge rebuild rates, the 5-second revocation promise, the 30-second interactive freshness budget, the 2-percentage-point recall gate, and the 75% chunk reuse rate — if the last of those is materially wrong, the fleet is undersized and the cost model is wrong by the same factor, which is why it is instrumented from the first day of the MVP.

## What is not here

The generative answer layer, the document stores themselves, the permission model, and the product user interfaces are all out of scope and belong to other teams. Multimodal corpora, learned chunking, multi-region index serving, lazy re-embedding of the dormant long tail, and persisting vectors as the durable artefact so a rebuild is IO-bound rather than GPU-bound are named in `ask.md` as Phase 3 and are deliberately absent from the views, so that nothing in the set reads as committed when it is not.
