Embedding Pipeline Service

Architecture Views

22 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

A shared internal platform that keeps vector embeddings and retrieval indexes in sync with a continuously edited document corpus, serving semantic search and ask-your-docs retrieval to fourteen product teams inside one collaboration-SaaS company. Read the set in order: the acts build an argument, and every view after act 2 exists to answer something an actor needed. One decision carries the whole design — an index is identified by its embedding contract and is immutable under it — and the views either honour that boundary or show what it costs.

Context and scope

What sits inside the boundary, what the platform depends on, and the shape of the path from an edit to a retrievable result.

People and journeys

Who the platform is for and what each of them gets to do. Three journeys get their own map: the one a knowledge worker runs daily, the one a product engineer runs once per corpus, and the one that is genuinely hard.
03 The people the product exists for Knowledge worker 3 M monthly Goal — I edited that page an hour ago. When I search for it, I want today's version — not the one from last week. Core journeys Find what I just wrote most-run journey Ask a question of the workspace Find documents like this one Tenant admin 25,000 organisations Goal — When I revoke access to a document, I need it gone from search immediately — not after whatever batch job runs next. Core journeys Revoke access to a document Erase a departing employee's content Prove what the platform holds The people who build on the platform Product engineer 14 consuming teams Goal — I want retrieval over my corpus without learning what a chunker is or owning a GPU node pool. Core journeys Ship a new corpus onboarding journey Measure my retrieval quality Choose a freshness tier Platform engineer on call for the fleet Goal — A better embedding model shipped this week. I want to adopt it without a quarter-long project or a quality regression nobody notices. Core journeys Upgrade the embedding model the hard one Rebuild a corrupted index Drain a backlog after an outage Privacy officer one per region Goal — I need to show that an erasure actually happened, in every store, including the ones nobody remembers exist. Core journeys Audit an erasure end to end Check a vector export control Machines in the cast Corpus change feed 8 M versions/day Goal — Hand over every edit once, in order per document, and be told plainly when I am being throttled. Core journeys Deliver an edit Deliver a tombstone Reconciliation sweep nightly per corpus Goal — Find the documents the feed never mentioned, before a user notices they are missing from search. Core journeys Compare source to ledger Confirm a suspected deletion Quality harness per corpus, nightly Goal — Say whether retrieval got worse, and which of the three causes the evidence actually supports. Core journeys Score the frozen query set Gate a contract cutover Who the Platform Is For, and What They Get to Do Person or role Journey / task Security / platform External / third party Three journeys get their own map: find what I just wrote, ship a new corpus, and upgrade the embedding model. v 1.0 · owner Data & AI Platform Architecture · date 2026-10 Actors and Their Core Journeys Five kinds of people and three machines, each with a goal in their own words and the journeys that goal turns into. HTML page SVG draw.io
05 Product engineer owns one product surface Goal — Get retrieval over my corpus into production without owning GPUs, a chunker or an index Trigger — A product commitment to ship search over a new content type this quarter Done when — Retrieval live, quality measured against a query set I wrote, and a cost figure I can defend 1 · Register self-service 2 · Backfill bulk lane 3 · Measure ◆ moment of truth 4 · Launch ◆ moment of truth 5 · Own it What they do Declares the corpus Picks a freshness tier Waits on the backfill Writes 80 eval queries Switches traffic on Watches the dashboard What the platform does Contract assigned Index provisioned Bulk lane, preemptible Progress + ETA Scores recall@10 Alias published Cost + freshness showback How it feels Confident Uneasy Blocked Where it hurts No ETA, unclear if it is stuck Nobody has labelled relevance Bill arrives without a cause What answers it Chunker chosen by corpus type Oldest-unembedded age as the ETA Harness scores an unlabelled set Recall gate before the alias flips Reuse rate + cache hit as the two cost ratios Journey — Ship a New Corpus The trough is measurement: nobody has ground truth on day one. The platform's answer is a frozen query set with reusable judgements, not a label budget. v 1.0 · owner Data & AI Platform Architecture · date 2026-10 Journey — Ship a New Corpus A product engineer onboarding retrieval over a new content type, and the day-one problem nobody has an answer to. HTML page SVG draw.io
06 Platform engineer on call for the fleet Goal — Adopt a better embedding model across 4.8 billion chunks without a quality regression or a quarter-long project Trigger — A new open-weights model benchmarks 6 points better on the corpus's own query set Done when — Every tenant served by the new contract, the old index retired, and a rollback that was never needed but stayed possible 1 · Evaluate on a sample 2 · Build beside dual-write 3 · Gate ◆ moment of truth 4 · Cut over ◆ moment of truth 5 · Retire What they do Pins a model digest Prices the migration Starts the shadow build Reads the recall diff Flips tenants in waves Reclaims the old index What the platform does New contract registered GPU hours estimated Both indexes live-fed Oldest unmigrated chunk Recall@10 vs outgoing Alias write per tenant 21-day rollback window How it feels In control Exposed Cornered Where it hurts Storage doubled for 3 weeks Scores not comparable across models A half-flipped estate during the wave What answers it Contract is identity, so the build is separate Dual-index overhead is a budgeted capacity state Gate on the frozen set, not on raw scores No query ever spans two contracts Rollback is an alias write, not a rebuild Journey — Upgrade the Embedding Model The trough is cut-over, and it is where the critical design decision earns itself: a per-tenant alias flip is safe only because a query never spans contracts. v 1.0 · owner Data & AI Platform Architecture · date 2026-10 Journey — Upgrade the Embedding Model The journey that justifies the whole architecture: adopting a better model across 4.8 billion chunks without a regression and without a project. HTML page SVG draw.io

Structure

The parts, their layers, and the interfaces across which they are reached — including the asymmetry that lets the control plane carry a lower availability target than retrieval.
08 Capture and preparation Ingest Connector workers per source type Admission quotas, lanes Change log Kafka, 13 mo Reconciliation sweep Argo CronWorkflow Prepare Extraction workers sandboxed, bounded Chunker versioned Normalised text MinIO Chunk ledger PostgreSQL, Patroni Embedding and index build Inference Batcher 3 lanes Embedding fleet KubeRay + TEI Vector cache Redis Reference probe rollout gate Index Index builder per contract Vector index Qdrant Lexical index OpenSearch Snapshots MinIO, RTO 4 h Serving Query path Retrieval gateway Envoy Query embedder same contract Hybrid fusion Citation resolver offsets Access control ACL filter fails closed Suppression list 5 s p99 Control plane Authority Contract registry PostgreSQL Index catalogue aliases in etcd Migration orchestrator Argo Workflows Quality and cost Quality harness ClickHouse Drift monitors Cost attribution showback Corpus sources 4 kinds Permission authority query-time Product surfaces 14 teams change feed retrieval can read? alias Platform Components — Container View Interface / broker Application we own Queue / topic Security / platform Data store External / third party event / async synchronous Omitted for clarity: the observability stack, secrets distribution, and the warehouse export. Each appears on its own view. v 1.0 · owner Data & AI Platform Architecture · date 2026-10 Platform Components The container view: four boundaries, the components inside each, and the three external systems they talk to. HTML page SVG draw.io

Data

What is stored, who owns it, what can be rebuilt and what cannot. The storage zones and the data model are where the critical design decision becomes a constraint rather than a principle.

Runtime

What actually happens when a document is edited, when a model is upgraded, and when a consumer asks a question — including the GPU economics that set the cost of all three.

Operations

Where it runs, what is watched, and the loop by which a contract is proposed, gated, served and retired.

Assurance

Why the platform is safe to put a tenant's corpus into, and what happens on each of the failures the requirement names.
22 What fails How it shows Immediate handling Recovery Residual risk Model endpoint drift Same name, new weights Reference probe diverges Rollout failed, not pipeline Pinned digest re-pulled Undetected if probe set is unrepresentative Partial document Some chunks fail Version never marked visible Previous version serves Retry, then quarantine A document stuck one version behind Change-feed gap Events never emitted Sweep finds discrepancy Re-ingest from source Ledger reconciled Up to one sweep interval blind Bad index build Low recall or corrupt Recall gate fails Alias never flips Rebuild or roll back Gate blind to a shift the frozen set misses Vector tier down Qdrant unavailable Retrieval errors spike Lexical fallback, flagged Snapshot restore, RTO 4 h Semantic recall lost while degraded Erasure not confirmed A store does not ack Unconfirmed erasure alert Suppressed on read path Retry to completion Bytes present though unreachable Backlog after an outage Hours of edits queued Oldest-unembedded age climbs Lane priority, bulk shed Surge GPU capacity Standard lane breaches while draining Assurance — Failure Classes and Their Handling Six more classes are named in the requirement and handled the same way. These seven are the ones that have changed a design decision. v 1.0 · owner Data & AI Platform Architecture · date 2026-10 Failure Classes and Their Handling Seven of the thirteen named failure classes, chosen because each one changed a design decision — with the residual risk stated. HTML page SVG draw.io

Embedding Pipeline Service — Architecture One-Pager

Keeping a vector representation of a continuously edited corpus accurate, fresh, permission-correct and affordable

An index is identified by its embedding contract and is immutable under it. Changing the contract builds a new index beside the old one and cuts over by alias; it never updates vectors in place.

Every workspace product people use daily has grown the same feature: a search box that understands what you meant, and an assistant that answers from your own documents with citations. The visible surface is a box. What decides whether it works is a pipeline nobody sees, keeping embeddings in step with a corpus that is edited eight million times a day, across twenty-five thousand organisations, under permissions the platform does not own. The hard part is not building it once. It is that embedding models improve every few months while a corpus lives for years — so a design in which a model change is a project will be two years out of date within two years, and the retrieval quality the entire product rests on will decay with no moment at which anyone decided to let it.

A shared internal platform. Product teams register a corpus and declare a freshness tier; the platform captures changes, extracts and chunks the text, embeds only what changed, builds immutable per-contract indexes behind mutable aliases, and serves hybrid retrieval with serve-time permission filtering to fourteen product surfaces. Open source throughout, self-hosted on Kubernetes: Kafka as the change log, PostgreSQL as the chunk ledger, MinIO for retained text and snapshots, Redis as the vector cache, a KubeRay GPU fleet for inference, Qdrant and OpenSearch for retrieval, etcd for aliases, Argo Workflows for rebuilds, ClickHouse for quality and drift.

What it is, and what it is not

An index identified by its embedding contract, immutable under itAn index as a long-lived mutable store that is upgraded in place, document by document
Chunk identity derived from content, so an edit costs what it changedChunk identity derived from position, so a paragraph insert re-embeds the document
The corpus as the truth, with every vector reproducible from retained text plus a pinned contractThe index as the truth, with vectors as precious state that has to be preserved
Permission resolved at serve time against the authority, failing closedPermission copied into the index at build time and refreshed when someone remembers
Revocation on the read path in seconds, indexing in minutesRevocation inheriting the pipeline's latency because it travels the same path
Cutover gated on measured retrieval qualityCutover gated on the build being finished

The decisions that are the architecture

01The embedding contract is the identity of an index, not configuration applied to one

A contract is (normaliser version, chunker version, model digest, pooling rule, dimensionality). A vector's primary key is (chunk hash, contract id), and the index builder refuses a vector whose contract does not match the index's. "We upgraded the model" therefore cannot mean a gradual, invisible, half-migrated result — the most common way these systems rot.

ADR-01

02Build beside, then flip

A new contract builds a shadow index that receives the live change stream alongside the serving one. Cutover is an alias write; rollback is the same write in reverse, against an index that is still built, still fed and still measured. The storage doubling and the dual-write window are budgeted capacity states with a declared 21-day life, not incidents.

ADR-02

03No query ever spans two contracts

Similarity scores from different models are not comparable and cannot honestly be normalised. Admitting that early is what makes progressive per-tenant cutover a real option and query-time fan-out across two models a trap, and it is why the alias is per tenant rather than global.

ADR-03

04The corpus is the truth; the chunk ledger is ours

Content and permissions stay with their owner. The platform's own system of record is the chunk ledger, and every index is a projection rebuildable from the ledger, the retained text and a pinned contract. That is what makes a corrupt index, a bad build and a wrong chunker one remedy instead of three disaster procedures.

ADR-04

05Chunk identity is a content hash, so an edit costs what it changed

A typical 41-chunk document edit produces three changed chunks. Everything downstream — GPU time, index writes, cost — is proportional to that number rather than to the document. Chunk reuse rate is therefore tracked as an operational signal, because it is the single number that sets the cost of the edit stream.

ADR-05

06The chunker version is a contract element

A chunker improvement is as much a full rebuild as a model change, and pretending otherwise produces an index whose chunk boundaries vary by when a document was last touched. Making the chunker a contract element prices chunking experiments honestly and keeps every stored vector reproducible.

ADR-06

07Three freshness lanes, with bulk on preemptible capacity

Interactive, standard and bulk, each with its own published batch wait. GPU inference is far cheaper per vector in large batches, and a large batch is bought with waiting — so freshness is spent deliberately, lane by lane, and a 14-day rebuild is affordable because bulk is preempted rather than interactive being queued behind it.

ADR-09

08A document version becomes visible atomically

The new version is retrievable only once every chunk of it is indexed; until then the previous version serves. A half-reindexed document matches on old and new wording at once and no reader can tell which they are seeing, which is strictly worse than being stale.

ADR-10

09Permission is resolved at serve time and fails closed

The platform asks the corpus's authority whether this subject may read this document, on every retrieval, and withholds everything when the authority does not answer. It is the one hard hot-path dependency in the design, accepted because the alternative — an ACL copy inside the index — is a second source of truth for the most consequential fact in the system.

ADR-13

10Revocation travels the read path, not the pipeline

A deletion or newly-restricted document stops being retrievable within five seconds through a suppression list consulted before the permission batch, while full erasure propagates through every store behind it. Revocation must not inherit the indexing latency it has nothing to do with.

ADR-14

11Model replicas are addressed by digest and gated by a frozen probe

A registry serving different weights behind the same name would silently invalidate every stored vector and present as "search got worse" months later. Every rollout embeds a frozen reference set and compares against stored expected vectors; divergence fails the rollout, not the pipeline.

ADR-15

12Cutover is gated on measured quality, and drift is monitored in three places

Recall@10 against a frozen evaluation set must not fall more than two percentage points against the outgoing contract. Corpus drift, model change and query drift all present as "search got worse" and only one is fixed by rolling back, so each is monitored separately and every alert says which the evidence supports.

ADR-16

13Snapshots of a rebuildable store exist for recovery time, not durability

A from-scratch rebuild takes three to fourteen days and satisfies no recovery-time objective anyone would sign, so indexes are snapshotted to reach a 4-hour RTO — and a restored snapshot is not trusted until the deletion log has been replayed over it, because otherwise a restore resurrects an erased document.

ADR-17

Why this should still hold in ten years

The decisions above are deliberately stated in terms of authority, identity and reproducibility rather than in terms of products. That is what should let them survive the thing most likely to change about this system, which is every component in it.

The boundary is not a technology

"An index is identified by its embedding contract and is immutable under it" says nothing about Qdrant, Kafka, PostgreSQL or Ray. A different vector store, a different log, a managed embedding API or an architecture not yet invented can each satisfy it, and everything downstream — dual-write, alias cutover, per-tenant progressive flip, rebuild as routine — follows from the boundary rather than from the stack.

It is designed for the one change that is certain

Embedding models will keep improving on a cadence of months. A design whose central operation is "adopt a new model" gets better as the field does; a design whose central operation is "serve the model we chose" gets worse. The cost the architecture pays — reproducibility, retained text, a rebuild budget — is the premium on that option, and it is paid once rather than per migration.

Reproducibility outlives any particular index

Because every vector is derivable from retained text plus a pinned contract, the platform can answer "what exactly produced this result" years later, and can move to a store that does not exist yet by rebuilding rather than migrating. The chunk ledger is the asset; the indexes are inventory.

The permission decision ages well

Resolving access at serve time against the owner's authority costs latency now and will keep costing it. But sharing models grow more complex over time, never less, and a platform that copied an ACL model in 2026 would be maintaining a divergent replica of a model nobody documented by 2030. Asking is slower and stays correct.

The failure posture does not depend on scale

Degrade a dimension rather than the service; declare staleness rather than imply freshness; fail closed on access and fail soft on everything else. These hold at a tenth of this scale and at ten times it, which matters because a platform's traffic is the assumption most likely to be wrong.

What will date

The specific numbers — 768 dimensions, 4,000 chunks a second, a 30-day text window, a 2-point recall gate — are all products of 2026 hardware and 2026 models and should be expected to move by an order of magnitude. They are stated as assumptions precisely so that moving them is an edit rather than an excavation. Quantisation, learned chunking and multi-vector retrieval are the three places where a change in the field would most plausibly force a structural revision rather than a numeric one.

Non-functional targets

Every target below is a stated assumption. They are included because a requirement without a number cannot be designed against, and because each of these drove at least one decision in the record.

QualityTargetHow it is metView
Retrieval availability ≥ 99.95% monthly on the retrieval path Stateless retrieval API across three zones; an index already built serves with the entire build path stopped 07
Retrieval latency p50 ≤ 45 ms, p99 ≤ 180 ms in-region for a hybrid top-50 over a tenant partition; p99 ≤ 400 ms above 50 M chunks Tenant-partitioned ANN index, query embedding on the same contract, batched ACL resolution 15
Freshness Interactive lane p50 ≤ 30 s, p95 ≤ 2 min, p99 ≤ 10 min; standard p95 ≤ 30 min; bulk none Lane admission at capture, per-lane batch wait, oldest-unembedded age as both SLI and scaling trigger 14
Revocation ≤ 5 s at p99 for a deleted or newly-restricted document to stop being retrievable Read-path suppression list checked before the permission batch, independent of indexing latency 10
Ingest availability ≥ 99.9% monthly, measured as change events durably logged and acknowledged Accept returns on durability in the change log; never waits on extraction or inference 13
Control-plane availability ≥ 99.5% monthly Deliberately lower: an outage freezes cutovers and widens freshness, and removes no retrieval 09
Throughput, live stream 3,000 chunks/s sustained, 5× burst for 15 minutes Autoscaled Ray Serve replicas on on-demand GPU capacity, lane-isolated 14
Throughput, rebuild 4,000 chunks/s background (≈ 14 days) or 18,500 chunks/s surge (≈ 72 h at ≈ 4.5× hourly cost) Bulk lane on preemptible GPU capacity, spilling to on-demand; fleet sized on this, not on 280/s mean 14
Durability, authoritative stores RPO 0, RTO 15 min for the chunk ledger, contract registry, index catalogue and deletion log PostgreSQL with Patroni, synchronous standby in a second zone; append-only deletion log 11
Recovery, indexes RTO 4 h from snapshot; a from-scratch rebuild is 3 to 14 days and satisfies no RTO Periodic index snapshots to MinIO, with the deletion log replayed over any restore before it serves 11
Quality gate Recall@10 must not fall more than 2 percentage points against the outgoing contract; quantisation ≤ 1 point Frozen per-corpus evaluation set scored by the quality harness; the gate blocks the alias flip 19
Cost ≤ $0.45 per million chunks embedded at background rate; ≤ $11 per million chunks per month to store and serve; full re-embed ≤ $2,200 GPU Chunk reuse rate and cache hit rate monitored as operational signals; migrations priced before they start 18

Scope

In scope

  • Corpus registration, change capture from feeds and webhooks, and a reconciliation sweep that finds what the feed missed
  • Extraction, OCR and versioned normalisation of the product's core formats, with offsets preserved back to the source document
  • Versioned chunking with content-hash chunk identity, and a chunk ledger that diffs document versions
  • Embedding inference on an operated GPU fleet addressed by model digest, with a vector cache and a rollout probe
  • Immutable per-contract vector indexes behind mutable aliases, with hybrid lexical retrieval and atomic version visibility
  • Serve-time permission filtering that fails closed, and a read-path suppression list giving 5-second revocation
  • Embedding contracts, dual-write migration, alias cutover and rollback, and migrations priced before they start
  • A frozen evaluation set per corpus, drift monitoring in three places, and per-tenant cost showback

Explicitly out of scope

  • The generative answer layer: prompting, synthesis, citation rendering and the conversation state behind them
  • The document stores, file stores and connected SaaS corpora themselves, which remain systems of record
  • The permission model: the platform resolves access against the owner's authority and never decides it
  • The product user interfaces, including any indexing-in-progress signal a surface may want to show
  • Choosing the embedding model, which arrives as a digest and becomes an input to a contract
  • Multimodal corpora, learned chunking, multi-region index serving and lazy long-tail re-embedding, all deferred to Phase 3

What to build first, and what it has to prove

The MVP is not a smaller version of the platform. It is the smallest system that can demonstrate the critical design decision is both true and affordable, because everything else in the package is downstream of that.

  1. One corpus, one tenant, change-feed ingestion with monotonic versions over a durable change log
  2. Extraction with offset mapping and retained normalised text; a versioned chunker with content-hash identity
  3. A chunk ledger that diffs document versions, with the reuse rate instrumented from day one
  4. An embedding fleet addressed by model digest, a vector cache, and the frozen reference probe on rollout
  5. One ANN index per tenant partition per contract, behind an alias, with atomic version visibility
  6. Serve-time permission filtering that fails closed, plus the read-path suppression list
  7. A second contract, built beside the first, dual-written, gated on recall@10, cut over and rolled back once on purpose
  • Edit a paragraph in a 300-chunk document. Prove that the number of chunks embedded is in single figures, and that the document is retrievable under its new text inside the interactive budget.
  • Revert that edit. Prove that no inference happened at all, because every chunk hash was already in the cache.
  • Register a second contract with a different model digest, build it beside the first, and cut over. Prove the flip is a single write and that retrieval never served results from both contracts at once.
  • Roll the cutover back. Prove it is the same single write, and that the previous index was still current because dual-write never stopped.
  • Revoke access to a document that is in the serving index. Prove it stops being retrievable inside five seconds, before any pipeline work has run.
  • Stop the entire build path — capture, extraction, inference — and prove retrieval still meets its p99 against the index already built.
  • Delete an index and rebuild it from the ledger and retained text alone, with no source system involved, and measure how long it took. That number is the migration budget.

Open risks, carried rather than hidden

RiskIf it landsResponse
Real-world chunk reuse on edit is materially below the assumed 75% The embedding fleet is undersized, the cost model is wrong by the same factor, and the interactive freshness budget is missed under normal load rather than under burst Instrument reuse rate per corpus from the first day of the MVP and treat it as a launch criterion, not a report. If a corpus's editing pattern rewrites whole documents, that corpus needs a different chunking strategy or a different freshness tier — and the number tells you which before a product has shipped on it.
The frozen evaluation set is unrepresentative, so the cutover gate passes a contract that is worse in production A migration completes, the alias flips, and retrieval quality falls for queries nobody wrote into the set. The rollback window is 21 days and the regression may not be noticed inside it Build the set from the real query log rather than from imagination, refresh a sampled slice of it each quarter, and pair the gate with implicit signals — answer acceptance, reformulation rate — watched for two weeks after each wave. Treat a gate pass with a falling acceptance rate as a failed migration.
The permission authority becomes the platform's availability ceiling Retrieval is contractually 99.95% but is in practice bounded by a system the platform does not own and cannot fail open on. Every authority incident is a retrieval incident Keep the ACL check batched and cacheable for the life of a request, publish the authority's own availability beside the retrieval SLO so the dependency is honest, and negotiate a target with its owners. Revisit pre-filtering only with measured recall loss in hand — an ACL copy trades an availability problem for a correctness one.
A learned or semantic chunker becomes clearly better, and chunking stops being deterministic The reproducibility principle weakens: a vector can no longer be derived from retained text plus a pinned contract unless the chunker model is itself pinned and retained like an embedding model Treat a learned chunker as a model from the outset — pinned by digest, probed on rollout, a contract element with the same rebuild cost. If that is unaffordable for a given corpus, the honest answer is that the corpus keeps a deterministic chunker, not that reproducibility is relaxed.
The dormant long tail never migrates, and the estate permanently spans contracts Storage for superseded indexes is never reclaimed, two or three contracts are maintained indefinitely, and the cost model that justified the architecture quietly inverts Make lazy re-embedding an explicit Phase 3 decision with a retirement deadline attached, not a drift. Measure the fraction of each corpus ever queried, and if the tail is large, dematerialise idle indexes rather than leaving them on an old contract — an index nobody queries should cost storage, not a maintained contract.

Architecture Decision Record — Embedding Pipeline Service

Why every component and every technology on these 22 views is what it is, and what each choice costs.

Seventeen decisions, grouped into six areas. Each carries the forcing question it answers, the context that made it necessary, how it is actually realised on a self-hosted Kubernetes stack, the alternatives including the ones that are right for a different organisation, the conditions that would flip it, why it should still hold as scale and technology change, and the transferable lesson. Read the one-pager first; it is the argument these records support.

Status of this document. Every quantity in this package is a stated assumption, sized for a collaboration-SaaS company with roughly 25,000 customer organisations, 3 million monthly active users, 400 million documents and 4.8 billion live chunks. None is measured from a production system. They are stated explicitly so they can be argued with and corrected: an explicit number forces the architecture to commit, and vagueness cannot be reviewed. Where a number sets a design boundary — the 14-day background rebuild, the 72-hour surge rebuild, the 5-second revocation, the 2-percentage-point recall gate — the record says so.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

Identity and immutability 4

The decisions that make a model change a routine, reversible operation instead of a project nobody can describe afterwards.

ADR-01The embedding contract is the identity of an index, and an index is immutable under it ADR-02A contract change builds a new index beside the old one, with dual-write, and cuts over by alias ADR-03No query is ever answered from two contracts, because similarity scores are not comparable across models ADR-04The corpus stays with its owner; the chunk ledger is the platform's own system of record

The cost of change 4

The decisions that make the price of an edit proportional to what changed, and the price of a rebuild knowable in advance.

ADR-05Chunk identity is a content hash of the normalised chunk text, not a position in the document ADR-06The chunker version is a contract element, so a chunking change is a full rebuild ADR-07A vector cache keyed by (chunk hash, contract), scoped to a tenant ADR-08Normalised text is retained for 30 days so a rebuild never touches the source systems

Freshness and correctness 4

The decisions about how fast the index follows the corpus, and what the platform refuses to show while it is catching up.

ADR-09Three freshness lanes, with bulk on preemptible capacity and the batch wait published per lane ADR-10A document version becomes retrievable atomically, or not at all ADR-11The change log is keyed by document, at-least-once, with a monotonic source version as the arbiter ADR-12A reconciliation sweep is a first-class component, and the absence of a document never implies a deletion

Permission and erasure 2

The decisions about the most consequential fact in the system — who may see what — and about making a deletion true in every store.

ADR-13Permissions are resolved at serve time against the owner's authority, and the platform fails closed ADR-14Revocation and deletion travel a read-path suppression list, not the indexing pipeline

Quality and drift 2

The decisions that let the platform say retrieval got worse, and say which of three indistinguishable causes the evidence supports.

ADR-15Model replicas are addressed by digest, and every rollout is gated by a frozen reference probe ADR-16Cutover is gated on a frozen evaluation set, and drift is monitored as three separate causes

Operating the fleet 1

The decisions about running GPUs, degrading honestly, and recovering a store that is formally disposable and practically irreplaceable in the moment.

ADR-17Indexes are snapshotted for recovery time, not durability, and a restore replays the deletion log

Technology by capability

Open source, self-hosted on Kubernetes, chosen deliberately and for two reasons that point the same way. The first is rotation: the five most recent use cases in this practice all went to a hyperscaler, and a sixth would teach that cloud's service catalogue rather than architecture. The second is that this topic is one where the managed option hides the lesson. Behind a per-token embedding API, the cost of a model migration is a line item on an invoice; on an owned GPU fleet and an owned index it is a capacity plan, a rebuild window, a spot-versus-on-demand decision and a storage doubling the architecture has to absorb. Every requirement in ask.md is written vendor-neutrally — "durable change log", not a product name — so the table below is one defensible realisation rather than part of the requirement.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Durable change log Kafka, partitioned by document identifier, 13-month retention via tiered storage to MinIO Open source, self-hosted Redpanda; NATS JetStream; Pulsar; a managed stream service Keyed partitioning gives per-document ordering, and long retention lets the transport double as the replay and deletion record ADR-11
Chunk ledger (system of record) PostgreSQL with Patroni, synchronous standby in a second zone, partitioned by tenant Open source, self-hosted CockroachDB; YugabyteDB; a managed relational service The one store with RPO 0, needing transactions, a cheap set-difference query for the diff, and operational familiarity ADR-04
Normalised text and snapshots MinIO, erasure-coded, objects under tenant-scoped keys Open source, self-hosted Ceph RGW; SeaweedFS; cloud object storage S3-compatible, cheap per TB, and the natural home for both retained text and index snapshots ADR-08
Vector cache Redis, partitioned per tenant, keyed by (chunk hash, contract) Open source, self-hosted Dragonfly; Valkey; Memcached A cost mechanism rather than a correctness one, so a simple, fast, losable store is exactly right ADR-07
Embedding inference KubeRay with Ray Serve, running a text-embedding server, replicas pinned by image and weight digest Open source, self-hosted KServe; Triton; vLLM-based serving; a managed embedding API Autoscaled replicas bound to node pools by taint, which is what makes three lanes with different capacity postures expressible ADR-15
Vector index Qdrant, one collection per (tenant partition, contract), snapshots to MinIO Open source, self-hosted Milvus; Weaviate; pgvector; Vespa; a managed vector service Cheap collection creation makes one-index-per-contract affordable, and payload filtering supports the version predicate atomic visibility needs ADR-01
Lexical index OpenSearch over the same chunk rows Open source, self-hosted Elasticsearch; Tantivy; Vespa as a single hybrid engine The degradation path when the vector tier is unavailable, and half of within-contract hybrid fusion ADR-03
Index aliases etcd, compare-and-swap per (tenant, corpus), watched by the retrieval gateway Open source, self-hosted Consul; the relational catalogue itself; a vector store's native alias feature A tiny, strongly consistent, watchable pointer store — exactly the shape of the alias, and separable from the catalogue's history ADR-02
Rebuild and migration orchestration Argo Workflows and CronWorkflows, holding migration state in PostgreSQL Open source, self-hosted Temporal; Airflow; Dagster Long-running, resumable, Kubernetes-native jobs for a 14-day rebuild, a nightly sweep and a restore ADR-02
Extraction and OCR Tika-based parsers and Tesseract in egress-restricted pods, bounded in time, memory, depth and expansion ratio Open source, self-hosted Unstructured; a commercial document-intelligence API The least-trusted compute in the system, so it must be isolated rather than merely sandboxed by convention ADR-08
Quality, drift and cost analytics ClickHouse for evaluation runs, drift series and cost attribution Open source, self-hosted DuckDB on object storage; Druid; a cloud warehouse Columnar scans over thirteen months of runs and series, which is a different access pattern from anything else in the platform ADR-16
Workload identity and secrets SPIRE for SVIDs, OpenBao for source credentials, Keycloak for human identity Open source, self-hosted cert-manager with an internal CA; HashiCorp Vault; a cloud identity service mTLS between every workload, short-lived source credentials rotatable without a pipeline outage, and operators separated from workloads ADR-13
Retrieval edge Envoy, terminating mTLS and resolving the alias from a watched cache Open source, self-hosted NGINX; Traefik; a service mesh ingress Resolving the alias at the edge from a cached pointer is what keeps a control-plane outage off the retrieval path ADR-02
Observability OpenTelemetry collection, Prometheus with Mimir, Grafana, Loki and Tempo Open source, self-hosted VictoriaMetrics; a managed observability platform Oldest-unembedded age per lane is both SLI and scaling trigger, so metrics have to be first-class rather than a side channel ADR-09

The decisions, and the alternatives that lost

Identity and immutabilityThe decisions that make a model change a routine, reversible operation instead of a project nobody can describe afterwards.

ADR-01

The embedding contract is the identity of an index, and an index is immutable under it

Accepted

When a better embedding model ships, what exactly happens to the 4.8 billion vectors already in the index?

Context
The obvious design treats the embedding model as configuration: a setting on the pipeline, pointing at a model name. It is simple, it needs no extra storage, and it has no answer to the question above. The honest answers available to it are all bad. Re-embed in place over weeks, and the index holds vectors from two models whose cosine similarities are being compared with each other — a correctness failure that produces no error and no alert, only slightly worse results. Re-embed in place over a weekend, and the cost is unbounded and the operation is unstoppable halfway. Or never re-embed, and the index is frozen on the model that happened to be current when the platform launched, while the field moves every few months. The same problem arrives from three other directions: a chunker improvement, a normaliser fix and a pooling change each invalidate stored vectors in exactly the same way, and a design that has no name for "the thing that produced this vector" cannot tell any of them apart.
Decision
An embedding contract is the tuple (normaliser version, chunker version, model digest, pooling and normalisation rule, dimensionality), and it is the identity of an index rather than configuration applied to one. Every stored vector records the contract that produced it; a vector's primary key is (chunk hash, contract id). The index builder refuses a vector whose contract does not match the index's. Changing any element of the contract therefore cannot modify an existing index — it can only produce a new one.
How it is realised on AWS
Contracts live in a PostgreSQL registry in the control plane, content-addressed by a hash of the tuple. The model digest is the registry digest of the served model image and weights, not a tag. Each tenant partition has one index per contract in Qdrant, named `<tenant>-<contract>-<n>`, and consumers never see that name: they resolve an alias held in etcd. The index builder reads the contract id from the catalogue when it opens a collection and rejects writes whose contract id differs, so the constraint is enforced at the only place vectors enter.
Options weighed
  • ChosenContract as index identity, indexes immutable under it: Makes every contract element — model, chunker, normaliser, pooling — a first-class, priced, reversible change with one mechanism
  • RejectedModel as pipeline configuration, re-embedding in place: Produces a heterogeneous index whose scores are meaningless across the heterogeneity, with no error and no way to describe the current state
  • RejectedModel version as a filterable field on the vector, one index for all contracts: Keeps one index at the price of making every query specify a model and every ANN structure hold incomparable neighbours
  • Right elsewherePin one model forever and never migrate: Right for a corpus that is loaded once and never improved on — a frozen archive, a regulatory snapshot. Wrong for a product whose search quality is the product
Consequences
What it buys
  • A model, chunker or normaliser change is one operation with one shape: build beside, measure, flip the alias, retire the old index
  • The state of the estate is always describable — every index names exactly what produced its contents, and every result can name the contract it came from
  • Rollback exists by construction, because the previous contract's index is a separate artefact rather than overwritten rows
  • Reproducibility is structural: given retained text and a contract id, any vector can be regenerated and compared
What it costs
  • Vectors must be reproducible, so the chunk ledger and the retained normalised text become load-bearing stores with their own RPO rather than conveniences
  • A full rebuild must fit a declared window and budget, which forces the fleet to be sized on the rebuild rate rather than the steady-state rate
  • Two indexes exist during every migration, so the storage floor includes a planned doubling
  • Nothing cheap is available for a small contract change: fixing a normaliser typo costs a full rebuild of the affected corpus
Choose differently when
If the corpus were static — loaded once, never edited, never re-embedded — the whole apparatus would be dead weight and a single mutable index would be correct. It would also flip if embedding models stopped improving, or if a model family emerged whose successive versions produced genuinely comparable vector spaces, which would make in-place upgrade safe. Neither looks likely inside this architecture's lifetime.
Why it holds up over time
The decision is stated in terms of identity and authority, not technology. It survives replacing Qdrant, replacing the GPU fleet with a managed embedding API, or moving to a vector representation that does not exist yet, because all of those are contract elements rather than architecture. What would date it is a change in the nature of embeddings themselves — multi-vector or late-interaction retrieval, where one chunk produces many vectors — and even then the contract concept extends rather than breaks.
LessonWhen a derived store is produced by something that will change, give the producer a name and make that name part of the store's identity. A pipeline whose output cannot say what made it has no migration story, only a series of incidents.
Shown on views06 12 16 19
ADR-02

A contract change builds a new index beside the old one, with dual-write, and cuts over by alias

Accepted

How does a new index become the serving index without a window in which retrieval is wrong, missing or inconsistent?

Context
ADR-01 establishes that a contract change produces a new index. That leaves the harder operational question of how it starts serving. The corpus does not stop being edited for the duration — at 8 million document versions a day, a 14-day rebuild sees over 100 million live edits land while it runs. An index built from a snapshot of the ledger and then switched to would be two weeks stale on the day it went live, and catching it up afterwards is the same problem again with less time. Meanwhile the old index must keep serving, correctly and freshly, throughout, because the migration is invisible to users and must stay that way.
Decision
A migration provisions the shadow index first, then enables dual-write, then begins re-embedding the ledger. From the moment dual-write is on, every live change is applied to both the serving and the shadow index under their own contracts. The shadow is therefore never stale on live edits, and "complete" means complete. Cutover is a write to the alias in the index catalogue. Rollback is the same write in reverse, and remains available for a declared 21-day retention period during which the superseded index continues to receive dual-write.
How it is realised on AWS
The migration orchestrator is an Argo Workflow holding the migration row in PostgreSQL. Enabling dual-write sets a flag the index-writer consumers read, after which each prepared chunk is embedded once per active contract and written to each matching index. Re-embedding walks the chunk ledger in tenant order through the bulk lane. Progress is published as the oldest unmigrated chunk age rather than a percentage, because a percentage hides a stalled tail. The alias flip is a single etcd compare-and-swap per tenant, and the retrieval gateway picks it up from a watch within a second.
Options weighed
  • ChosenShadow index with dual-write, atomic alias flip per tenant: The shadow is current on live edits throughout, so cutover is instantaneous and rollback is symmetric
  • RejectedBuild from a ledger snapshot, then catch up, then flip: The catch-up is the same race with less margin, and the index is at its most stale exactly when it starts serving
  • RejectedQuery-time fan-out across both contracts with score normalisation: Requires comparing scores across models, which ADR-03 establishes is not honestly possible; also doubles query cost permanently
  • Right elsewhereBlue/green at the cluster level rather than the index level: Reasonable where one tenant occupies a cluster. Here it would force all 25,000 organisations to flip together, removing progressive rollout
Consequences
What it buys
  • Cutover is a single write, so there is no window in which retrieval is partially migrated
  • Rollback is the same single write against an index that is still built, still fed and still measured — not an emergency path exercised once under pressure
  • Progressive per-tenant flip becomes possible, which turns a 25,000-organisation migration into a series of small, observable ones
  • The migration can be abandoned at any point by stopping dual-write and retiring the shadow, leaving no partial state
What it costs
  • Peak storage during the overlap is roughly 1.6× steady state (60% additional, as the superseded index is retained 21 days)
  • Every prepared chunk is embedded once per active contract while dual-write is on, which raises live-stream GPU cost for the window
  • The orchestrator holds real state — which tenants are flipped, what dual-write is on — and becomes a component that must itself be correct
  • A long migration means a long period in which the estate spans contracts, invalidating cross-tenant quality comparisons
Choose differently when
With a corpus small enough to rebuild inside a maintenance window — say under 50 million chunks at surge rate, a few hours — the whole dual-write apparatus is unnecessary and a stop-the-world rebuild is simpler and cheaper. The decision is justified by a rebuild measured in days against a corpus that is edited continuously.
Why it holds up over time
Build-beside-and-flip is one of the oldest patterns in operations and has outlived every technology it has been implemented on. What is specific here is dual-write during the overlap, and that follows from the corpus being live rather than from any component. It would only weaken if rebuild time fell to the point where staleness during a snapshot build stopped mattering.
LessonA cutover you can reverse with the same action that performed it is a cutover you can afford to attempt. Build the reverse path first and the forward path becomes routine.
Shown on views06 16 19
ADR-03

No query is ever answered from two contracts, because similarity scores are not comparable across models

Accepted

During a migration, can a query search both the old and the new index and merge the results?

Context
It is the first idea everyone has, and it looks like it solves the whole problem: if the new index is only 40% built, search both and merge, and the user gets the best available answer throughout. The difficulty is that the numbers being merged are not measurements of the same thing. Two embedding models place text in different spaces with different geometries; a cosine similarity of 0.82 in one is not stronger, weaker or equal to 0.82 in the other, and the relationship between them is not a monotonic function that can be calibrated away. Score normalisation across models — min-max over the candidate set, z-scoring, rank fusion — produces a number that looks comparable and is not, and the resulting ordering is an artefact of the normalisation rather than of relevance. Worse, the failure is silent and partial: most results are plausible, and the ones that are wrong look like ordinary retrieval noise.
Decision
A single query is resolved against exactly one contract. The alias a tenant resolves points at one index; there is no fan-out and no cross-contract merge. A tenant is either on the old contract or the new one, never both. Reciprocal-rank fusion is used within a contract to combine vector and lexical retrieval, where the two signals are deliberately different in kind and the fusion is over ranks rather than scores.
How it is realised on AWS
The alias in etcd maps (tenant, corpus) to exactly one index id. The retrieval gateway resolves it once per request and passes the contract id to the query embedder, so the query is embedded by the same model that produced the index. A contract mismatch between the query embedder and the index is a hard rejection rather than a degraded answer, and the rejection is counted.
Options weighed
  • ChosenOne contract per query, progressive per-tenant cutover: Keeps every score comparison inside one space, and makes the migration unit a tenant rather than the estate
  • RejectedFan-out across contracts with score normalisation: Produces an ordering that is an artefact of the normalisation, failing silently and partially — the worst failure mode available
  • RejectedFan-out with rank fusion instead of score normalisation: Avoids the score problem and still merges two partial views of the corpus, so recall depends on migration progress in a way nobody can reason about
  • Right elsewhereFan-out during migration only, accepted as temporary: Defensible where retrieval is advisory and a user can see both result sets labelled by source. Not defensible behind an assistant that cites one answer
Consequences
What it buys
  • Every score the platform compares lives in one space, so ranking means what it appears to mean
  • Progressive per-tenant cutover becomes the natural migration shape, since the tenant is already the unit of index addressing
  • A query embedder on the wrong contract is a loud rejection rather than a quiet quality regression
  • Within-contract hybrid fusion stays legitimate, because rank fusion over two deliberately different signals is a different claim from score normalisation over two models
What it costs
  • A tenant mid-migration sees results from the old contract until its flip, so the new model's benefit arrives in waves rather than gradually
  • Cross-tenant comparisons of retrieval quality are invalid while the estate spans contracts, which complicates exactly the measurement a migration most wants
  • The alias becomes per tenant rather than global, multiplying the number of pointers the catalogue holds
Choose differently when
If a model family shipped successive versions with provably aligned vector spaces — or if a cheap, reliable projection between two models' spaces existed and was itself evaluated — fan-out would become safe and the constraint would relax. Research in that direction exists; nothing production-ready does. It would also flip for a product where retrieval is advisory and the user can see which index a result came from.
Why it holds up over time
The underlying fact is mathematical rather than technological: two independently trained embedding spaces have no canonical alignment. That will not change because a vendor ships a new model. The decision should therefore outlast every component in this architecture, and its main risk is that someone re-proposes fan-out each time a migration feels slow.
LessonBefore merging two sets of numbers, ask whether they are measurements of the same thing. If they are not, a normalisation step does not fix it — it hides it, which is worse.
Shown on views06 15 16
ADR-04

The corpus stays with its owner; the chunk ledger is the platform's own system of record

Accepted

When the index and the corpus disagree about what a document says, which one is wrong — and what does the platform own well enough to rebuild from?

Context
A retrieval platform sits in an awkward position: it must hold a derived representation of content it does not own, under permissions it does not control, and it must be able to recover from its own failures without asking 400 million documents to be re-delivered. Two tempting designs fail. Treating the index as authoritative makes a corrupt shard, a bad build or a wrong chunker unrecoverable except by full re-ingestion from source — a load spike the corpus owners will refuse. Copying the corpus wholesale into the platform makes recovery easy and creates a second system of record for content and permissions, with all the divergence, residency and erasure obligations that implies.
Decision
Content and permissions remain with their owners; the platform holds a pointer, a source version and a derived representation. The platform's own system of record is the chunk ledger: one row per (document, version, chunk hash) with offsets, structural path, contract version, vector reference and embedding status. Every vector index is a projection rebuildable from the chunk ledger, the retained normalised text and a pinned contract, with no input from any source system for documents inside the text retention window.
How it is realised on AWS
The chunk ledger is PostgreSQL with Patroni, synchronous standby in a second zone, RPO 0 and RTO 15 minutes — the only hard durability promise in the design. Normalised text sits in MinIO keyed by (document, version, normaliser version) with a 30-day expiry carried as a column rather than enforced by a cleanup job. A rebuild is an Argo Workflow that walks the ledger, reads text from MinIO, embeds through the bulk lane and writes a fresh index. Nothing but the ledger, the contract registry and the deletion log is backed up as truth.
Options weighed
  • ChosenCorpus with its owner, chunk ledger as the platform's record, indexes as projections: Recovers from any index failure without touching source systems, and keeps exactly one system of record for content
  • RejectedThe index as the system of record: Every index failure becomes a full re-ingestion from source — a load spike the corpus owners will not accept, and an RTO measured in days
  • RejectedFull corpus copy inside the platform: Easy recovery at the price of a second source of truth for content and permissions, with the residency and erasure obligations that follow
  • Right elsewhereNo retained text; always re-fetch from source on rebuild: Right where the corpus is small or the source is cheap to read. At 400 million documents it makes every rebuild a source-system incident
Consequences
What it buys
  • A corrupt shard, a bad build, a wrong chunker and a failed migration all have one remedy: rebuild from the ledger
  • The platform's backup surface is small — a relational ledger, a registry and a log — rather than 12 TB of index data
  • Erasure has a single authoritative place to start from, and reconciliation has something to compare against
  • The derived representation can be thrown away deliberately, which is what makes ADR-01 affordable
What it costs
  • The chunk ledger is a single-region relational store holding billions of rows and must be operated accordingly, with partitioning and vacuum pressure as real concerns
  • Retained text is an additional store with its own retention policy, residency obligation and cost
  • Outside the 30-day window, a rebuild does hit source systems, so the retention period is a bet about rebuild cadence
  • Permissions resolved rather than held means a hot-path dependency on a system the platform does not own (ADR-13)
Choose differently when
If the corpus owner offered a durable, cheap, high-throughput read path with no load concern — or if the platform and the corpus were the same service — the retained-text store would be redundant and re-fetching on rebuild would be simpler. If the platform were instead the corpus owner, a single store would serve both purposes and this record would collapse into a schema decision.
Why it holds up over time
"The derived store is disposable and the ledger is not" is a statement about authority. It survives changing the ledger's technology, the text store's technology and the index's technology. The 30-day window is the part that will date, and it is stated as an assumption for that reason.
LessonOwn the smallest thing you cannot recreate, and make everything else recreatable from it. The size of what you must back up is a design output, not a fact of life.
Shown on views01 11 12

The cost of changeThe decisions that make the price of an edit proportional to what changed, and the price of a rebuild knowable in advance.

ADR-05

Chunk identity is a content hash of the normalised chunk text, not a position in the document

Accepted

When someone edits one paragraph of a 300-chunk document, how many chunks does the platform have to embed?

Context
This single question decides whether the platform is affordable. At 8 million document versions a day, the difference between re-embedding every chunk of an edited document and re-embedding only the changed ones is roughly a factor of four in GPU spend and the difference between meeting and missing the interactive freshness budget. The naive identity — document plus ordinal — makes a paragraph inserted near the top shift every subsequent chunk's ordinal, so a one-word change invalidates the document. A byte-offset range has the same problem, shifted. The alternative, hashing the normalised chunk text, makes identity independent of position: unchanged text is the same chunk wherever it moved to. It brings its own complications, principally that the natural place to anchor a citation is a position rather than a hash, and that identical text appearing in thousands of documents collapses to one identity.
Decision
A chunk's identity is a hash of its normalised chunk text. Position — ordinal, offsets, structural path — is carried as attributes of the chunk row rather than as its key. The ledger diff between two document versions is therefore a set difference over hashes, and only hashes absent from the previous version enter the embedding stage. Chunk reuse rate is reported per corpus as a first-class operational metric.
How it is realised on AWS
The chunker emits, for each chunk, a SHA-256 over the normalised chunk text plus the chunker version. The ledger holds (chunk_hash PK, document_id, version, ordinal, offset_start, heading_path), so a chunk that moved keeps its hash and gets a new ordinal. The diff is a single query against the previous version's hash set. The vector cache (ADR-07) is keyed by the same hash, which is why a reverted edit costs nothing at all. Deduplication is scoped per tenant, so identical boilerplate collapses within an organisation but never across one.
Options weighed
  • ChosenContent hash of the normalised chunk text: Makes the cost of an edit proportional to what changed, and makes reverted edits and shared templates free
  • RejectedDocument plus ordinal: An insertion near the top invalidates every chunk below it, so a one-word edit costs a whole document
  • RejectedDocument plus byte-offset range: Same failure as the ordinal, and additionally fragile to any normaliser change that shifts offsets
  • Right elsewhereStructural address (heading path plus index within section): Attractive where citations must be permanently stable and documents are structurally rigid — a legal corpus, a standards library. Costs reuse on exactly the edits users make most
Consequences
What it buys
  • A typical 41-chunk document edit produces about three changed chunks, making the edit stream affordable at 280 embeddings per second mean rather than roughly 1,100
  • A reverted edit, a duplicated page and a shared template all resolve to cached vectors and cost no inference
  • The diff is a cheap set operation rather than a text comparison, so preparation stays CPU-bound and predictable
  • Reuse rate becomes a measurable, per-corpus number that predicts cost before a product ships on it
What it costs
  • A citation cannot be anchored on the chunk key, so offsets must be carried separately and kept correct through normalisation — which is why ADR-08 retains the text that produced them
  • Cross-document deduplication is deliberately limited to a tenant, giving up hit rate to avoid an existence side channel across an organisation boundary
  • A chunk shared by many documents complicates deletion: the vector may only be dropped when the last referring document is gone
  • A normaliser change alters every hash in the corpus, which is why the normaliser is a contract element and not a tweak
Choose differently when
For a corpus that is written once and never edited — ingested archives, published records — reuse is irrelevant and a positional identity would be simpler and would make citation anchoring trivial. The decision is justified entirely by continuous editing, and its value scales directly with edit rate.
Why it holds up over time
Content addressing is a durable idea and does not depend on the hash function, the store or the model. The chunker-version component of the hash is what keeps it honest as chunking evolves. What could date it is a move to multi-vector retrieval, where one chunk yields several vectors and identity has to extend to the vector within the chunk.
LessonThe identity you choose for a derived unit decides the cost of every future change to its source. Pick it by asking what a typical edit should cost, not by asking what is easiest to compute.
Shown on views10 12 13
ADR-06

The chunker version is a contract element, so a chunking change is a full rebuild

Accepted

Is improving the chunker a tuning change or a migration?

Context
Chunking is the most-tinkered-with part of a retrieval system and the least respected. Teams change chunk size, overlap, boundary rules and context prefixes casually, because each change is a few lines and the effect on quality is immediately visible on a handful of queries. The consequence is almost never considered: a chunker change alters chunk boundaries, so it alters chunk text, so it alters every chunk hash, so every stored vector for that corpus is for text that no longer exists as a chunk. An index that has quietly accumulated three chunking regimes — depending on when each document was last edited — retrieves fragments of three different shapes and cannot be reasoned about. The temptation is to treat chunking as cheap because it needs no GPU; the reality is that a chunking change costs exactly as much as a model change, because both invalidate the same vectors.
Decision
The chunker version is an element of the embedding contract, with the same consequences as the model digest. A chunker change produces a new contract and therefore a new index, built beside and cut over by alias. Chunkers are selected per corpus type by the platform, not by the consuming team, and a chunker change is priced and gated exactly as a model change is.
How it is realised on AWS
The chunker is a versioned service; its version string enters the contract tuple and the chunk hash. Chunker configuration — size, overlap, boundary rules, context prefix composition — is part of the version rather than runtime configuration, so there is no path by which a deployed chunker behaves differently from the one that built the serving index. A chunker experiment runs by registering a candidate contract against one corpus and scoring it on that corpus's frozen evaluation set before any migration is proposed.
Options weighed
  • ChosenChunker version as a contract element: Prices chunking honestly and keeps every stored vector reproducible from retained text
  • RejectedChunker as runtime configuration: Produces an index with several chunking regimes depending on last-edit date, and breaks reproducibility silently
  • RejectedChunker version stored per chunk, mixed within one index: Makes the heterogeneity visible without making it harmless: retrieval still compares fragments of different shapes
  • Right elsewhereConsumer-chosen chunker per corpus: Right where consumers are sophisticated and corpora are few. Here it would hand fourteen teams a lever whose cost is a full rebuild they do not pay for
Consequences
What it buys
  • Retrieval over any index is retrieval over chunks of one shape, which is a precondition for the scores meaning anything
  • A chunker improvement is evaluated on a frozen query set before it is adopted, rather than shipped because it looked better on three examples
  • Reproducibility holds: retained text plus a contract regenerates any vector exactly
  • Chunking experiments become cheap at corpus scale and expensive at estate scale, which is the correct incentive
What it costs
  • Chunking improvements are rare, because each is a 14-day background rebuild and a 21-day storage overlap for the affected corpus
  • Consuming teams lose a lever they will expect to have, and the platform owes them an evaluation path instead
  • The chunker becomes a release-managed component with its own versioning discipline, which is heavier than a config flag
  • Normalisation and chunking are coupled through the hash, so a normaliser fix also triggers a rebuild
Choose differently when
If re-embedding were nearly free — a model small enough to run on CPU at corpus scale, or cached vectors that survive a chunking change, which they cannot — chunking could reasonably be runtime configuration. It also flips for a corpus small enough that a rebuild is minutes, where the discipline costs more than it saves.
Why it holds up over time
The coupling between chunk boundaries and stored vectors is inherent, not technological, so this holds as long as retrieval is over embedded chunks. It would be genuinely revisited by a retrieval architecture that embeds documents whole, or by late-interaction models where chunking is less determinative.
LessonFind the cheap-looking change that invalidates expensive derived state, and give it the same ceremony as the expensive change. The accounting, not the diff size, decides what a change costs.
Shown on views04 12 19
ADR-07

A vector cache keyed by (chunk hash, contract), scoped to a tenant

Accepted

The same chunk text appears in a reverted edit, a copied page and a shared template. How many times is it embedded?

Context
ADR-05 makes chunk identity content-derived, which immediately raises the possibility of never embedding the same text twice. The opportunity is substantial: in a collaboration corpus, template boilerplate, copied meeting notes, repeated headers and reverted edits mean a meaningful share of chunk hashes recur. The complication is scope. A cache shared across all tenants maximises hit rate and introduces an existence side channel — a caller who can observe a cache hit learns that some other organisation holds that exact text, which for a document fragment is a real disclosure. A cache scoped per tenant gives up cross-tenant hits and closes the channel.
Decision
A vector cache is maintained, keyed by (chunk content hash, contract id) and partitioned per tenant. A cache hit skips inference entirely. The cache is explicitly a cost mechanism rather than a correctness mechanism: it may be cold, partial or lost without affecting what the platform returns, only what it spends.
How it is realised on AWS
Redis, partitioned by tenant, holding the vector bytes against the composite key with a TTL long enough to cover an edit-and-revert cycle and a bulk re-run. The batcher consults the cache before forming a batch, so a fully cached document never reaches the GPU. Cache hit rate is a first-class operational signal alongside chunk reuse rate, because together they explain almost all variation in cost.
Options weighed
  • ChosenCache keyed by (chunk hash, contract), tenant-partitioned: Captures the within-tenant recurrence that dominates, with no cross-tenant inference possible
  • RejectedGlobally shared cache across all tenants: Higher hit rate, and an observable existence side channel on exact document text across an organisation boundary
  • RejectedNo cache; rely on the ledger diff alone: Loses the free cases entirely — a reverted edit would cost full inference, which is the most common edit pattern after a save
  • Right elsewhereCache in the index rather than separately (read-before-write): Sensible where the index is cheap to query by key. Here it puts build-path load on the serving tier, which is the one thing with a latency SLO
Consequences
What it buys
  • A reverted edit costs nothing at all, which matters because save-then-undo is a normal editing pattern rather than an edge case
  • Templates and duplicated pages within an organisation embed once, lowering the effective cost of large collaborative corpora
  • Bulk re-runs after a transient failure are cheap, because the work already done is still cached
  • Because the cache is a cost mechanism only, it can be resized, flushed or lost during an incident without any correctness review
What it costs
  • Cross-tenant recurrence — which at 25,000 organisations is probably large — is deliberately given up
  • Another stateful component to operate, size and monitor, holding derived data subject to the same tenant key controls as a vector anywhere else
  • Hit rate becomes a cost dependency: a cache incident shows up as a GPU bill rather than as an outage, which is harder to notice
  • TTL tuning is a real decision — too short loses the revert case, too long holds vectors for contracts being retired
Choose differently when
In a single-tenant deployment, or with a tenant population small and mutually trusting enough that content-existence disclosure is acceptable, a global cache would be correct and materially cheaper. The decision is driven by multi-tenancy, not by economics.
Why it holds up over time
Content-addressed caching of expensive derived values is stable. The tenant-scoping decision is a privacy judgement that should be revisited only with an explicit disclosure analysis, and the pressure to revisit it will grow as cost does — which is the main reason the reasoning is recorded here.
LessonA deduplication key that spans a tenant boundary is a disclosure channel, however innocuous the value looks. Decide the scope of a cache on who could learn something from a hit, not only on what it saves.
Shown on views10 14 18
ADR-08

Normalised text is retained for 30 days so a rebuild never touches the source systems

Accepted

When the platform rebuilds an index, where does the text come from?

Context
ADR-04 makes indexes rebuildable, and ADR-01 makes rebuilds routine rather than exceptional. That raises a question with an unattractive default answer: a rebuild that reads from source systems is a rebuild that hammers 400 million documents out of services that exist to serve users, plus a re-run of OCR on every scanned file. At 18,500 chunks a second surge rate, the source read load during a migration would dwarf normal product traffic. Corpus owners would quite reasonably refuse. Retaining the extracted text solves it, at the cost of holding a second copy of tenant content — with the residency, encryption and erasure obligations that implies, and a storage bill proportional to the corpus.
Decision
Normalised text is retained in object storage keyed by (document, version, normaliser version), for 30 days, with the expiry carried as data rather than enforced by a job. A rebuild reads retained text where it is inside the window and re-fetches from source only where it is not. The retained copy is treated with exactly the sensitivity of the source document, under tenant-scoped keys, and is erased on a deletion like any other store.
How it is realised on AWS
MinIO with erasure coding, objects encrypted under tenant-scoped keys, with `expires_at` on the corresponding ledger row so retention is visible and queryable rather than implicit in a lifecycle rule. A rebuild workflow partitions the ledger, reads text objects in bulk, and only falls back to the connector fetch path for documents whose text has expired — which it reports as a distinct count, because that count is the source-system load a migration will cause.
Options weighed
  • ChosenRetain normalised text for 30 days; re-fetch only outside the window: Makes re-chunking and re-embedding free of source load for recently touched documents, which is most of the working corpus
  • RejectedRetain nothing; always re-fetch and re-extract on rebuild: Turns every migration into a source-system incident and re-runs OCR at corpus scale for no benefit
  • RejectedRetain indefinitely: Removes the fallback path entirely, at the cost of a permanent full second copy of all tenant content and its obligations
  • Right elsewhereRetain raw bytes rather than normalised text: Right if extraction itself is expected to change often. Here the normaliser is a contract element changed rarely, and raw bytes are larger and no more useful
Consequences
What it buys
  • A chunker or model change re-enters the pipeline at the chunk stage, so a rebuild is GPU-bound and predictable rather than bounded by somebody else's API
  • OCR is paid once per document version rather than once per rebuild, which for a scanned corpus is the dominant preparation cost
  • The 30-day window covers the documents users actually touch, so in practice most of a rebuild reads from storage
  • Expiry as data means retention is auditable and a privacy officer can query it rather than inferring it from a bucket policy
What it costs
  • A second copy of tenant content exists, with residency, encryption, access-control and erasure obligations identical to the original
  • Storage cost proportional to 30 days of document versions, which for an actively edited corpus is not small
  • Rebuilds of cold documents still hit source systems, so the migration load profile depends on how old the corpus's tail is
  • The window is a bet: a corpus whose rebuild cadence is longer than 30 days gets little benefit from it
Choose differently when
If extraction became nearly free and source systems offered a bulk, throttle-tolerant export path, retaining nothing would be cleaner and would remove a whole class of privacy obligation. Conversely, if migrations became frequent enough that the cold tail dominated, the window would need to grow — which is a cost decision, not an architectural one.
Why it holds up over time
The principle — keep the expensive intermediate, not the cheap one — is stable. The 30-day figure is explicitly an assumption and should be expected to move with observed rebuild cadence and with the cost ratio between storage and extraction.
LessonWhen a derived store must be rebuildable, decide early which intermediate you keep. Rebuildability that depends on someone else's API is a promise you do not control.
Shown on views10 11 12

Freshness and correctnessThe decisions about how fast the index follows the corpus, and what the platform refuses to show while it is catching up.

ADR-09

Three freshness lanes, with bulk on preemptible capacity and the batch wait published per lane

Accepted

GPU inference is far cheaper in large batches, and a large batch is bought with waiting. Who pays the wait?

Context
The economics here are stark and frequently ignored. Embedding throughput per GPU rises steeply with batch size up to the point of saturation, so the difference between embedding chunks as they arrive and embedding them in large batches is a large multiple in cost per vector. Waiting is therefore the cheapest lever the platform has, and freshness is what it spends. Running one undifferentiated queue forces a single answer for everyone: either every edit waits, in which case the interactive promise is unmeetable, or nothing waits, in which case a 4.8-billion-chunk rebuild is unaffordable. Worse, one queue lets a single tenant's 400,000-document backfill sit in front of another tenant's edit from ten seconds ago.
Decision
Work is classified into three lanes at admission — interactive, standard and bulk — each with its own batch size, maximum wait and freshness commitment. The interactive lane has a short wait and accepts poor batch efficiency. The bulk lane has no freshness SLO, runs only on preemptible capacity, and is preempted rather than allowing interactive work to queue behind it. Each lane's maximum wait is a published number, so the freshness-for-cost trade is explicit rather than a tuning accident.
How it is realised on AWS
Admission assigns a lane from the corpus's declared freshness tier and the work's origin — a live change event is interactive or standard, a backfill or re-embedding run is always bulk. Each lane is a Kafka consumer group with its own batcher configuration. Ray Serve holds separate deployments bound to node pools by taint: interactive and standard on on-demand GPU nodes, bulk on spot nodes that spill to on-demand only when idle on-demand capacity exists. Oldest-unembedded age is tracked per lane and is both the freshness SLI and the autoscaling trigger.
Options weighed
  • ChosenThree lanes with published per-lane batch wait, bulk on preemptible capacity: Lets freshness be bought where it is perceived and cost be saved where it is not, with a 14-day rebuild made affordable by spot
  • RejectedOne queue, batch for cost: Makes the interactive promise unmeetable and lets one tenant's backfill delay another tenant's edit
  • RejectedOne queue, no batching: Meets freshness and makes a full rebuild cost several times its budget, which breaks the decision the whole architecture rests on
  • Right elsewherePer-tenant queues rather than per-lane: Stronger isolation, and at 25,000 tenants it fragments batches until batch efficiency collapses — right at hundreds of tenants, wrong at tens of thousands
Consequences
What it buys
  • The interactive promise (p50 30 s) and the rebuild budget (14 days, ~$2,200 GPU) are simultaneously achievable, which no single-queue design manages
  • A tenant's bulk backfill cannot delay another tenant's live edit, which is the isolation property that actually gets complained about
  • Spot capacity withdrawal lengthens a migration and affects nothing else, so the cheapest capacity carries the most interruptible work
  • Oldest-unembedded age per lane distinguishes a large backlog from a stuck one, which queue depth cannot
What it costs
  • Three sets of batcher parameters and three node-pool postures to operate, tune and reason about
  • Interactive batches are small and therefore expensive per vector; the architecture pays a premium precisely where users notice
  • A migration's duration becomes dependent on spot availability, making completion estimates probabilistic
  • Draining a large backlog deliberately breaches the standard lane to protect the interactive one — the right ordering, and still a breach
Choose differently when
If embedding inference became cheap enough that batch size stopped mattering — a much smaller model, or hardware where the batch curve flattens early — lanes would be unnecessary complexity and one queue would be correct. The whole structure exists to monetise waiting, and it disappears if waiting stops being worth anything.
Why it holds up over time
The batch-size-versus-latency curve is a property of accelerator hardware and has held across several generations. The specific rates will move by orders of magnitude; the shape of the trade, and therefore the need to assign the wait deliberately, should not.
LessonWhen a resource is much cheaper in bulk, the question is not whether to batch but who pays the wait. Answer it per class of work, publish the answer, and the trade stops being an accident.
Shown on views05 14 17 18
ADR-10

A document version becomes retrievable atomically, or not at all

Accepted

While three of a document's chunks are being embedded, what does retrieval return for that document?

Context
The convenient answer is: whatever is in the index. Chunks are independent rows, each upserted as its vector arrives, and retrieval simply sees a mixture — most chunks from version 40, the three changed ones from version 41. It costs nothing to implement and it produces the worst failure mode this platform has. A reader searching for text they just removed finds it, because the old chunk is still there. A reader searching for text they just added finds it, in a document whose surrounding chunks contradict it. The assistant cites two passages from the same document that disagree. None of this raises an error, and none of it is attributable: the document looks wrong in a way that reads as the model being bad. A stale document, by contrast, is a single coherent artefact that is simply behind — and the platform already has a mechanism for saying so.
Decision
A document version is retrievable only once every chunk belonging to that version is indexed. Until then, retrieval continues to serve the previous complete version. Visibility is a property recorded on the document version in the ledger, and the index excludes chunks of versions not yet marked visible.
How it is realised on AWS
The ledger's document_version row carries `visible_at`, set only when every chunk hash for that version has a vector in the target index. Chunks are written to the index tagged with their version, and retrieval filters on the visible version per document. Superseded chunks are removed in the same transition, so there is no window in which two versions are both visible. A version that cannot complete — a poison chunk, a repeated inference failure — never becomes visible and the document stays one version behind, which is reported as a distinct metric.
Options weighed
  • ChosenAtomic per-document-version visibility: A document is always one coherent version, and staleness is declarable rather than a silent mixture
  • RejectedPer-chunk visibility as vectors land: Produces self-contradicting documents with no error, which reads to users as the retrieval being bad
  • RejectedAtomic across the whole corpus rather than per document: Would make freshness for any document hostage to the slowest document in the corpus
  • Right elsewherePer-chunk visibility with a version filter applied by the consumer: Workable where there is one sophisticated consumer. With fourteen teams it moves a correctness invariant into fourteen codebases
Consequences
What it buys
  • Retrieval over any document returns one internally consistent version, which is the precondition for a citation meaning anything
  • Staleness becomes a single declarable fact per document — the version and its age — rather than an unknowable mixture
  • A stuck document is visible as a document stuck at a version, which is a diagnosable condition rather than a quality complaint
  • Deleted text actually stops matching at the moment the new version becomes visible
What it costs
  • Freshness for a document is bounded by its slowest chunk, so a 310-chunk p99 document is slower than its median chunk suggests
  • The index must support a version filter cheaply, which constrains the choice of vector store and adds a predicate to every query
  • A poison chunk freezes a document at its previous version indefinitely until the quarantine is resolved
  • The ledger becomes part of the visibility path, coupling retrieval correctness to a store that is otherwise build-path only
Choose differently when
If documents were small enough that a version was always one or two chunks, atomicity would be free and the decision trivial. It would also flip for a corpus of append-only records, where an older chunk can never contradict a newer one and a mixture is harmless.
Why it holds up over time
The underlying argument — a partially updated document is worse than a stale one — is about human interpretation, not technology, and does not weaken with scale. The implementation detail most likely to change is where the version filter is applied as vector stores grow richer predicate support.
LessonDecide what the unit of consistency is before deciding how fast to update it. Speed applied to the wrong unit produces artefacts nobody can diagnose.
Shown on views04 10 13
ADR-11

The change log is keyed by document, at-least-once, with a monotonic source version as the arbiter

Accepted

Two edits to the same document arrive out of order. Which one ends up in the index?

Context
Change feeds from real systems are at-least-once, occasionally out of order, and sometimes replayed in bulk after an incident at the source. The platform cannot ask for exactly-once delivery because no source offers it honestly. If order within a document is not guaranteed, the older edit can win and the index silently holds text the user replaced — the most expensive kind of wrong, because nobody notices until a search returns deleted wording. The platform also needs the ability to replay the log, both to recover from its own pipeline failures and to reprocess a corpus whose preparation was wrong, which means the log has to be a durable record rather than a transport.
Decision
Accepted changes are appended to a durable, replayable change log partitioned by document identifier, so all changes to one document are processed in order by one consumer. Every change carries a source-assigned monotonic version, and the pipeline refuses to apply a version lower than the one already recorded in the ledger. Delivery is assumed at-least-once; idempotence comes from the ledger's keys, not from the transport's promises.
How it is realised on AWS
Kafka with the document identifier as the partition key, 13-month retention through tiered storage to MinIO, which makes the log both the transport and the deletion-and-change record the compliance story depends on. The ledger holds the applied source version per document; a consumer that reads a lower version acknowledges and drops it, counting the drop. Replay is an Argo Workflow reading from an offset or from the archive.
Options weighed
  • ChosenDocument-keyed durable log, at-least-once, monotonic version arbiter: Per-document ordering where it matters, idempotence in the data model, and replay for free
  • RejectedTrust the source feed's ordering and delivery: No source offers this honestly, and the failure is a document silently holding replaced text
  • RejectedTimestamp rather than source version as the arbiter: Clock skew between source and platform makes the arbiter non-monotonic exactly when two edits are close together
  • Right elsewhereGlobal ordering across the whole corpus: Correct where cross-document transactions matter. Here it would serialise 8 million versions a day for a guarantee nothing needs
Consequences
What it buys
  • Two rapid edits to one document cannot race, so the index converges on the newest version without a reconciliation pass
  • A bulk replay from the source is harmless: every already-applied version is dropped idempotently
  • The log doubles as the replayable record that makes a pipeline fix reprocessable and a deletion auditable
  • No cross-document ordering is promised, which keeps throughput linear in partitions
What it costs
  • A hot document — a page being collaboratively edited by twenty people — is a hot partition, and its throughput is one consumer's
  • The platform depends on the source supplying a monotonic version; where one does not exist, the connector must synthesise one and the guarantee is only as good as that
  • 13-month log retention is a storage commitment made for the compliance story rather than for the transport
  • Nothing guarantees cross-document consistency, which occasionally surprises a consumer expecting a moved page and its parent to update together
Choose differently when
If every source offered a durable, ordered, replayable change feed of its own with a long retention, the platform's log would be a redundant copy and consuming the source feed directly would be simpler. In practice the log also serves the compliance and replay requirements, so it would have to be replaced by something rather than removed.
Why it holds up over time
At-least-once plus idempotence in the data model is the one delivery design that has survived every generation of messaging technology, precisely because it assumes the weakest guarantee. The monotonic-version arbiter is equally stable; what will change is the broker.
LessonDo not buy ordering you do not need, and do not accept the absence of ordering you do. Partition by the unit whose internal order actually matters, and get idempotence from your own keys.
Shown on views10 13 22
ADR-12

A reconciliation sweep is a first-class component, and the absence of a document never implies a deletion

Accepted

A document was edited and no change event ever arrived. How does the platform find out?

Context
Change feeds miss things. A webhook fails and is not retried, a connector is restarted mid-batch, a source deploys a bug, a tenant re-authorises an integration and the first hour is lost. Each of these is individually rare and collectively certain across 25,000 organisations and four source kinds. The failure is silent and permanent: the document sits in the index at an old version forever, and the only detector is a user noticing that search does not find something they wrote. The obvious fix — periodically compare the source's document inventory against the ledger — introduces its own hazard, because the same comparison that finds a missed edit also appears to find deletions. A source API that is partially unavailable, paginating badly or filtering by a permission the connector lost will under-report its own inventory, and a sweep that treats absence as deletion will then erase a tenant's corpus from the index.
Decision
A reconciliation sweep runs periodically per corpus, comparing the source's document inventory and versions against the chunk ledger. A document present in the source at a newer version than the ledger is re-ingested. A document absent from the source is never deleted on that basis: it is a candidate that must be confirmed by an explicit per-document check against the source, and a sweep whose inventory is implausibly smaller than the previous sweep's aborts rather than acting.
How it is realised on AWS
An Argo CronWorkflow per corpus, nightly by default and configurable per freshness tier. It pages the source inventory, diffs against the ledger by (document id, source version), and enqueues re-ingestion in the standard lane. Deletion candidates are checked individually through the connector's get-document path; a 404 confirms, anything else does not. A sweep aborts if the inventory count has fallen by more than a declared fraction since the last successful sweep, and the abort pages a human.
Options weighed
  • ChosenPeriodic sweep for missed edits, deletion only on explicit confirmation: Catches feed gaps without giving a flaky source API the power to erase a tenant's retrievable corpus
  • RejectedNo sweep; trust the feed: Guarantees a slowly growing population of permanently stale documents, discovered only by users
  • RejectedSweep that treats absence as deletion: Converts any source-side pagination bug or permission change into mass removal from the index
  • Right elsewhereFull periodic re-ingestion instead of a diff: Simpler and correct for a small corpus. At 400 million documents it is a continuous rebuild of everything
Consequences
What it buys
  • A missed edit is corrected within one sweep interval rather than never, which bounds the worst case for freshness
  • A flaky or partially-unavailable source cannot cause removal, because removal requires a positive confirmation
  • Sweep discrepancy count becomes a direct measure of each source connector's reliability, which is otherwise invisible
  • The abort threshold turns a catastrophic class of bug into a page rather than an incident
What it costs
  • The sweep is real load on source inventory APIs, which has to be negotiated and rate-limited per corpus
  • Up to one sweep interval of blindness remains, so the sweep is a bound on the failure rather than a fix for it
  • Deletion confirmation is per document and therefore slow, so a genuine bulk deletion propagates through tombstones rather than through the sweep
  • Another scheduled component whose own failure is silent unless its successful-run recency is itself monitored
Choose differently when
If a source offered a verifiable, gap-free change feed with a sequence number the platform could check for continuity, the sweep would be replaceable by a continuity check — far cheaper and strictly better. It would also change shape for a corpus small enough to re-inventory continuously.
Why it holds up over time
"Verify what you were told, and never infer destruction from silence" is a durable operational principle. The sweep's implementation will change with whatever the sources offer; the prohibition on treating absence as deletion should not.
LessonEvery event-driven integration needs a reconciliation path, and that path must be unable to cause the destructive action. Detection and deletion are different privileges.
Shown on views03 10 22

Permission and erasureThe decisions about the most consequential fact in the system — who may see what — and about making a deletion true in every store.

ADR-13

Permissions are resolved at serve time against the owner's authority, and the platform fails closed

Accepted

Who decides whether this user may see this chunk, and what happens when that decision cannot be obtained?

Context
This is the most consequential fact in the system and the one most often got wrong by copying. The tempting design stores an ACL alongside each vector and filters inside the index: one round trip, excellent recall, and no hot-path dependency. It fails on staleness. Sharing changes constantly in a collaboration product — a document moved between spaces, a group's membership edited, a guest removed — and an ACL copied at index time is wrong from the moment it is written. Worse, the copy is a reimplementation of someone else's permission model, including inheritance, group nesting, link sharing and per-space defaults, maintained by a team that does not own the semantics. Divergence is not a risk; it is the steady state. The alternative — ask the authority on every retrieval — is slower and introduces a dependency the platform cannot fail open on, because failing open means showing a user a document they were explicitly denied.
Decision
The platform stores an opaque reference to a document's permission state and resolves access against the corpus owner's authority at query time, in a batch, for the candidate set. Any failure of that resolution — timeout, error, ambiguity — excludes the affected chunks and the response reports that results were withheld. The platform never decides access itself and never caches an allow decision beyond the life of a request.
How it is realised on AWS
The retrieval gateway collects the distinct document identifiers in the candidate set, consults the read-path suppression list (ADR-14) first, then issues one batched authorisation call for the remainder, with a strict timeout. A partial response excludes the documents not positively allowed. The withheld count is returned to the consumer and recorded; the audit record carries counts, never content. Allow decisions are held only for the duration of the request.
Options weighed
  • ChosenServe-time batched resolution against the owner's authority, fail closed: Always current and never a second source of truth for the fact that matters most
  • RejectedACLs copied into the index, filtered pre-search: Fast and stale: a reimplementation of someone else's permission semantics that diverges as its normal state
  • RejectedPhysical index partitions per principal group: Removes the hot-path dependency and generates an unbounded number of partitions under a real sharing model
  • Right elsewhereServe-time resolution with a short-lived decision cache: Right where revocation latency can be measured in minutes. Here the tenant admin's five-second promise forbids caching an allow
Consequences
What it buys
  • Access is always evaluated against the current state of a model the platform does not own and does not have to track
  • A revocation takes effect at the next query, with no index write and no propagation delay to reason about
  • Fail-closed means the worst outcome of a dependency failure is a missing result, never a disclosure
  • Withheld counts make the behaviour legible to a consumer, instead of results silently thinning
What it costs
  • The authority is a hard hot-path dependency and therefore the real ceiling on retrieval availability, whatever the platform's own SLO says
  • Post-filtering requires over-fetching candidates, and where a user may read a small fraction of a tenant's corpus, recall suffers measurably
  • Batched authorisation adds latency to every query and makes retrieval p99 partly somebody else's number
  • An authority incident presents as retrieval returning little or nothing, which is a confusing symptom for a product team to triage
Choose differently when
If the corpus owner published a change feed of permission events with a bounded propagation guarantee, a pushed permission cache would become defensible and would remove the hot-path dependency — this is the single most valuable thing a corpus owner could offer the platform. It also flips for a corpus where sharing is effectively static, such as a published handbook.
Why it holds up over time
Sharing models grow more complex over time, never less. A platform that copied a permission model in 2026 would be maintaining a divergent replica of an undocumented model by 2030. Asking is slower and stays correct, which is why this should outlast the components around it.
LessonNever hold a second copy of the most consequential fact in your system. If the copy would be stale the moment you write it, resolve instead — and make the failure of resolution exclude rather than permit.
Shown on views15 20 21
ADR-14

Revocation and deletion travel a read-path suppression list, not the indexing pipeline

Accepted

A tenant admin revokes access to a document. How long until it stops being retrievable?

Context
If revocation travels the same path as an edit, it inherits that path's latency: minutes at p50, ten minutes at p99, longer during a backlog. For an edit that is acceptable and honestly declarable. For a revocation it is not: the whole point of revoking access is that it takes effect now, and "the search index will catch up within ten minutes" is not an answer a privacy officer or an administrator will accept. The deeper problem is that erasure across every store — ledger, retained text, vector cache, several index versions, snapshots — is genuinely a distributed operation that cannot be instantaneous. So the platform needs two different promises: one about the read path, which can be fast, and one about erasure, which must be complete.
Decision
Deletion and permission revocation are written immediately to a suppression list consulted on the read path, before the permission batch, giving a 5-second p99 promise that the document stops being retrievable. Full erasure — chunk ledger, retained text, vector cache, every index version and any snapshot restored later — proceeds behind that, retried to completion, with tamper-evident evidence that it happened and an alert when a store has not confirmed inside the SLO.
How it is realised on AWS
Suppression is a small, replicated, read-optimised structure the retrieval gateway consults in-process with a short refresh interval; the write is synchronous on the admin path. Erasure is an Argo Workflow fanning out to each store, recording per-store confirmation in the deletion log. Any index snapshot restored later has the deletion log replayed over it before it serves traffic, which is what stops a restore resurrecting an erased document.
Options weighed
  • ChosenRead-path suppression for immediacy, asynchronous erasure for completeness: Two promises for two different obligations, neither compromised by the other
  • RejectedRevocation through the indexing pipeline: Inherits the pipeline's latency and its backlog, which is the one thing a revocation must not do
  • RejectedSynchronous erasure across every store before acknowledging: Makes the admin action as slow and as failure-prone as the slowest store, and offers no answer for a later snapshot restore
  • Right elsewhereRely on serve-time ACL resolution alone, with no suppression list: Defensible when the authority's own propagation is fast and a deletion is always visible to it. Here a deleted document may vanish from the authority too, and absence is not a denial
Consequences
What it buys
  • Revocation is 5 seconds at p99 while indexing stays at minutes, and neither promise constrains the other
  • Erasure is auditable per store rather than assumed, so "we deleted it" is evidence rather than a claim
  • A snapshot restore cannot resurrect an erased document, which closes the most commonly missed erasure hole
  • A store that fails to confirm raises an alert while the read path is already correct, so the incident is not a disclosure
What it costs
  • A second mechanism on the read path, which must itself be highly available: suppression unavailable means fail closed, which is correct and still an outage
  • The suppression list grows and needs its own compaction once erasure is confirmed across all stores
  • Two promises to explain to consumers and auditors, and the distinction is easy to lose in a status page
  • Bytes may remain present-but-unreachable while a retry is outstanding, which is correct on the read path and imperfect on disk
Choose differently when
If erasure across every store could be made genuinely fast — a far smaller corpus, or stores that all support cheap targeted deletion — one synchronous path would be simpler and the suppression list would be unnecessary complexity. The split exists because erasure at this scale cannot be fast and revocation must be.
Why it holds up over time
The separation of "stop showing it" from "remove every copy" is a permanent consequence of holding derived data in several stores, and regulatory expectations have moved consistently towards requiring both rather than one. This should hold.
LessonWhen one action implies two obligations with incompatible time constants, build two mechanisms. Collapsing them means the fast promise is as slow as the slow one, or the slow one is quietly not kept.
Shown on views10 20 21

Quality and driftThe decisions that let the platform say retrieval got worse, and say which of three indistinguishable causes the evidence supports.

ADR-15

Model replicas are addressed by digest, and every rollout is gated by a frozen reference probe

Accepted

How would the platform know that the model serving requests today is not the one that built the index?

Context
This is the quietest catastrophic failure available to a retrieval system. A model is referenced by a name or a tag; a registry rebuild, a vendor update, a mutable tag pushed over, a different quantisation applied at serving time, or a kernel version that changes numerics slightly — and the vectors produced today are no longer in the same space as the vectors in the index. Nothing errors. Similarities drift by amounts too small to notice per query and large enough to degrade ranking across millions of them. The symptom arrives weeks later as "search feels worse", at which point the cause is indistinguishable from corpus drift, query drift, or a chunking change nobody logged. By then the index is an unknown mixture and the only remedy is a full rebuild.
Decision
Models are addressed by digest — of the image and the weights — never by a name or a tag, and the digest is an element of the embedding contract. Every model rollout embeds a frozen reference set of chunks and compares the output vectors against stored expected values. Divergence beyond a declared tolerance fails the rollout; it never fails the pipeline, which continues on the previous replicas.
How it is realised on AWS
The serving image and weights are pinned by content digest in the Ray Serve deployment, and the contract registry records that digest. A reference set of a few thousand chunks, with their expected vectors, is stored alongside the contract. On every replica start the probe runs before the replica is marked ready; a replica whose vectors diverge never receives traffic. The tolerance accounts for legitimate non-determinism in floating-point accumulation and nothing more.
Options weighed
  • ChosenDigest addressing plus a frozen reference probe on every rollout: Turns an invisible, delayed, unattributable failure into a failed deployment at the moment it is introduced
  • RejectedModel name or tag with a version label: A mutable reference to the thing whose identity the entire index depends on
  • RejectedDigest addressing without a probe: Catches a changed artefact and not a changed numeric environment — a different GPU, driver or kernel can move vectors with the same digest
  • Right elsewherePeriodic probing rather than rollout gating: Better than nothing and detects the problem after some vectors are already wrong. Worth adding alongside, not instead
Consequences
What it buys
  • A silently changed model becomes a failed rollout rather than a slow quality regression nobody can attribute
  • The index's contents are genuinely identified by the contract, which is what makes ADR-01's guarantee real rather than nominal
  • The probe also catches hardware and driver changes that move numerics without changing any artefact
  • A rollout failure degrades nothing: the previous replicas keep serving, so the safe outcome is also the default one
What it costs
  • Every rollout costs a probe before readiness, adding latency to deployments and a small amount of GPU time
  • A tolerance has to be chosen, and choosing it badly produces either flapping rollouts or a probe that detects nothing
  • The reference set must be representative; an unrepresentative set passes a model that moved in the part of the space that matters
  • Model upgrades can no longer be done by moving a tag, which removes a convenience operators will miss
Choose differently when
Nothing plausible makes this unnecessary while the index depends on a model's exact numerics. If embeddings became robust to small numeric differences — or if retrieval moved to a representation where small shifts provably did not affect ranking — the probe could relax. Digest addressing would still be right.
Why it holds up over time
Content addressing of the artefact that produced persistent derived state is as durable a practice as exists in software. The probe is the part that adapts, since what needs checking depends on how much of the numeric environment the platform controls.
LessonIf a derived store depends on a computation's exact output, pin the computation by content and verify it on every deploy. A mutable reference to the thing your data depends on is a silent corruption waiting for a convenient moment.
Shown on views14 19 22
ADR-16

Cutover is gated on a frozen evaluation set, and drift is monitored as three separate causes

Accepted

Retrieval got worse. Did the corpus change, did the model change, or did the questions change?

Context
All three produce the same complaint and the same dashboards, and only one of them is fixed by rolling back. A corpus that has drifted in subject matter — a company that moved into a new product line — makes old queries match worse because the relevant content is genuinely different. A model that changed makes everything match differently. A query mix that moved away from what the corpus covers makes retrieval look worse while the index is unchanged and correct. Treating all three as one metric means every alert triggers the same investigation and the same instinct, which is to roll back the last contract — correct a third of the time. The second difficulty is that none of this can rest on labelled relevance data, because nobody will fund labelling at corpus scale and a gate that needs labels will be skipped at the first deadline.
Decision
Each corpus maintains a frozen evaluation set of queries with recorded expected results, scored for recall@10 and MRR on a schedule and on every contract change. A contract cutover is gated on recall@10 not falling more than two percentage points against the outgoing contract; completeness alone never opens the gate. Separately, the distribution of stored vectors and the distribution of incoming queries are monitored independently, and every quality alert states which of the three causes the evidence supports, or that it cannot tell.
How it is realised on AWS
Evaluation sets are built from the real query log rather than from imagination, with judgements recorded once and reused across contracts. The quality harness runs in ClickHouse-backed batch jobs, scoring the set against a named index and storing the run against (contract, corpus). Vector distribution — norm, centroid, pairwise-similarity spread — and query distribution are tracked as separate time series with a 3σ alarm against a 7-day rolling baseline. Implicit signals (result selection, answer acceptance, reformulation, abandonment) are collected under a declared telemetry contract and attributed to a contract version.
Options weighed
  • ChosenFrozen evaluation set as the gate, three drift signals monitored separately: Gives a cutover an objective veto and makes a quality alert diagnosable rather than merely alarming
  • RejectedGate on completeness; monitor quality afterwards: Allows a fully built, measurably worse contract to start serving, with the regression found inside the rollback window only by luck
  • RejectedGate on implicit signals alone: Implicit signals lag by days and are confounded by product changes, so they cannot gate an operation that happens in seconds
  • Right elsewhereGate on human-labelled relevance judgements per migration: The strongest gate and the right one where a labelling budget exists. Here it would be skipped under deadline, which is worse than a weaker gate that is always run
Consequences
What it buys
  • A migration can be refused on evidence, by a mechanism rather than by an argument
  • Judgements are bought once and reused on every later contract, so the cost of the gate falls over time
  • Separating the three drift causes means an alert points at an action instead of at a rollback reflex
  • Thirteen months of evaluation history makes a slow regression across two consecutive migrations visible rather than lost at each cutover
What it costs
  • The gate is exactly as good as the evaluation set, and a set drawn from an old query mix will certify a contract that is worse on current traffic
  • Maintaining a frozen set per corpus is real work for fourteen consuming teams, and the platform owes them a harness to make it cheap
  • Drift monitoring on 4.8 billion vectors is sampled, so a localised shift can be missed
  • Implicit signals require a telemetry contract with consumers, which is a privacy and governance conversation as much as an engineering one
Choose differently when
If a reliable offline proxy for retrieval quality emerged — a model-based judge trustworthy enough to gate on — the frozen set would become a regression guard rather than the primary gate, and migrations would get faster. Conversely, if a corpus genuinely had no stable query distribution, the frozen set would stop being meaningful and the gate would have to be implicit-signal based with a slow, staged rollout.
Why it holds up over time
The three-causes distinction is a property of the problem and will not change. The gate mechanism should be expected to improve as evaluation methods do, and the architecture deliberately places the gate at a single point — the alias flip — so that improving it does not touch anything else.
LessonWhen several different causes produce one symptom, instrument the causes separately before instrumenting the symptom. And prefer a weaker gate that is always run to a stronger one that is skipped under deadline.
Shown on views06 18 19

Operating the fleetThe decisions about running GPUs, degrading honestly, and recovering a store that is formally disposable and practically irreplaceable in the moment.

ADR-17

Indexes are snapshotted for recovery time, not durability, and a restore replays the deletion log

Accepted

The vector index is formally disposable. Why back it up — and what makes a restore safe?

Context
ADR-04 establishes that every index is a projection rebuildable from the ledger and retained text. The tempting conclusion is that indexes need no backup at all: if one is lost, rebuild it. The arithmetic refuses. A full rebuild runs 14 days at the background rate and 72 hours at surge, and no product owner will accept a 72-hour retrieval outage for a single-shard loss, let alone 14 days. So the rebuildability argument, which is correct about durability, says nothing useful about recovery time. Snapshots close that gap — and introduce a hazard of their own, because a snapshot taken before an erasure, restored after it, puts the erased document back into a serving index. That is a compliance failure produced by a recovery procedure, which is the kind that gets missed because the people who design restores and the people who design erasure are rarely the same people.
Decision
Vector indexes are snapshotted periodically to object storage, giving a 4-hour RTO, and the snapshots are explicitly a recovery-time mechanism rather than a durability claim — the ledger remains the record. Any restored snapshot has the deletion log replayed over it before it is allowed to serve traffic, and that replay is part of the restore procedure rather than a follow-up task.
How it is realised on AWS
Qdrant snapshots to MinIO on a schedule per collection, retained for a short window since their value decays. Restore is an Argo Workflow: fetch the snapshot, load it into a collection marked not-serving, replay the deletion log from the snapshot's timestamp forward, verify the suppression list against the collection, then make it serving by alias write. A restore that cannot complete the replay does not proceed; the fallback is a rebuild, slower and always correct.
Options weighed
  • ChosenSnapshots for RTO, deletion-log replay before serving: A 4-hour recovery that cannot resurrect erased content, with rebuild as the always-correct fallback
  • RejectedNo snapshots; rebuild from the ledger on any loss: Correct about durability and offers a 72-hour-to-14-day RTO, which no consumer will accept
  • RejectedSnapshots restored directly, with erasure reconciled afterwards: Serves erased documents for the length of the reconciliation, which is a disclosure created by the recovery procedure
  • Right elsewhereContinuous replication of the index instead of snapshots: Better RTO and right where the index is the record. Here it replicates a projection continuously and still needs the deletion replay on promotion
Consequences
What it buys
  • Recovery time and durability are addressed by separate mechanisms, each honest about what it provides
  • A restore cannot resurrect an erased document, because the replay is inside the procedure rather than after it
  • Snapshot retention can be short, since the ledger is the record and an old snapshot has little value
  • A failed replay degrades to a rebuild, so the unsafe path is never the fast path
What it costs
  • Snapshot storage for 12 TB of index data, for a mechanism that exists only to shorten recovery
  • The restore procedure is more complex than loading a snapshot and must be rehearsed, or it will be wrong when needed
  • The deletion log becomes load-bearing for recovery as well as compliance, raising its own availability requirement
  • A snapshot is a full copy of tenant-derived content and inherits every encryption, residency and access obligation
Choose differently when
If rebuild time fell inside an acceptable RTO — a much faster fleet, or vectors persisted in a columnar store so that a rebuild is IO-bound rather than GPU-bound, which is the Phase 3 option on the roadmap — snapshots would become unnecessary and this record would be superseded. That is the single change that would most simplify the platform's recovery story.
Why it holds up over time
The distinction between durability and recovery time is permanent and routinely conflated. The deletion-replay requirement will only strengthen as erasure obligations do. What may date is the need for snapshots at all, if rebuild becomes cheap.
Lesson"It is rebuildable" answers a durability question and says nothing about recovery time. Check which question your backup policy is actually answering — and make sure the restore path cannot undo a deletion.
Shown on views11 17 22

Every package used, in one table

The terms this package uses in a specific sense. Where a more common word exists and was rejected, the reason is given — usually because the common word implies something the architecture deliberately does not promise.

PackageWhat it isWhat it does hereConsidered instead
Embedding contract The tuple (normaliser version, chunker version, model digest, pooling and normalisation rule, dimensionality). The identity of an index and of every vector in it. The unit that is proposed, priced, built, gated, served and retired. "Model version", which omits the chunker and normaliser and so understates what invalidates a vector.
Chunk hash A hash over the normalised chunk text together with the chunker version. Chunk identity, the key of the ledger diff, and the key of the vector cache. "Chunk id", which suggests an assigned identifier and hides that identity is derived from content.
Chunk ledger One relational row per (document, version, chunk hash) with offsets, structural path, contract version, vector reference and status. The platform's own system of record, the basis of every reconciliation, and the input to every rebuild. "Metadata store", which implies something incidental rather than the one store with RPO 0.
Alias A mutable pointer from (tenant, corpus) to exactly one index id, with the previous id retained. The only thing consumers resolve, and the mechanism of both cutover and rollback. "Index name", which would couple every consumer to the contract currently in service.
Dual-write The state in which every prepared chunk is embedded once per active contract and written to each matching index. What keeps a shadow index current on live edits, so "complete" means complete and cutover is instantaneous. "Backfill", which describes only the historical half and omits the live half that makes cutover safe.
Freshness lane One of interactive, standard or bulk: a class of work with its own batch size, maximum wait, capacity posture and freshness commitment. Where the batch-efficiency-for-freshness trade is made explicit and assigned. "Priority", which implies ordering only and hides that each lane has a different capacity posture and a different promise.
Atomic version visibility A document version is retrievable only once every chunk of that version is indexed. What prevents a document matching on both its old and its new wording at once. "Eventual consistency", which is true of the platform and says nothing about the unit over which consistency is guaranteed.
Suppression list A read-path structure naming documents that must not be returned, consulted before the permission batch. The 5-second revocation promise, held separate from the minutes-long indexing path and from the slower erasure fan-out. "Deny list", which suggests a policy artefact rather than a time-critical correctness mechanism.
Reference probe A frozen set of chunks with stored expected vectors, embedded on every model rollout and compared. The detector for a model that changed behind an unchanged name, or a numeric environment that moved. "Smoke test", which implies checking that the service answers rather than that it answers identically.
Drift Three distinct conditions with one symptom: the corpus moved, the model moved, or the query distribution moved. Monitored as three separate signals, because only one of the three is fixed by rolling back a contract. Using "drift" for all three, which is how a quality alert becomes a rollback reflex that is right a third of the time.
Chunk reuse rate The fraction of a document's chunks that are unchanged between two consecutive versions. With cache hit rate, the number that sets the cost of the edit stream and sizes the embedding fleet. "Cache hit rate" alone, which measures the second-order saving and misses the first-order one.
Rebuild Reconstructing an index from the chunk ledger, the retained normalised text and a pinned contract, with no source-system input. The single remedy for a corrupt index, a bad build, a wrong chunker and a failed migration. "Restore", which implies returning to a previous state rather than deriving the current one.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.