# Architecture Decision Record — Embedding Pipeline Service

*Embedding Pipeline Service · Solution Architecture v1.0 · open source, self-hosted on Kubernetes · Data & AI Platform Architecture · 2026-10 · 22 views · 17 architecture decision records*

The argument these decisions serve is summarised in the [Architecture One-Pager](architecture-one-pager).

Seventeen decisions, grouped into six areas. Each carries the forcing question it answers, the context that made it necessary, how it is actually realised on a self-hosted Kubernetes stack, the alternatives including the ones that are right for a different organisation, the conditions that would flip it, why it should still hold as scale and technology change, and the transferable lesson. Read the one-pager first; it is the argument these records support.

> **Status of this document.** Every quantity in this package is a stated assumption, sized for a collaboration-SaaS company with roughly 25,000 customer organisations, 3 million monthly active users, 400 million documents and 4.8 billion live chunks. None is measured from a production system. They are stated explicitly so they can be argued with and corrected: an explicit number forces the architecture to commit, and vagueness cannot be reviewed. Where a number sets a design boundary — the 14-day background rebuild, the 72-hour surge rebuild, the 5-second revocation, the 2-percentage-point recall gate — the record says so.

## How to read a record

- **Question:** The forcing question: why a decision was needed at all.
- **Context:** The requirement, the scale and the constraint that make it hard.
- **Decision:** What this architecture does, stated so it can be checked.
- **How it is realised on AWS:** The concrete mechanism: which service or package, configured how, in which subscription.
- **Options weighed:** Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- **Consequences:** What the choice buys and what it costs, both kept visible.
- **Choose differently when:** The conditions that would flip the decision for your system.
- **Why it holds up over time:** What keeps the decision right as scale, staff and technology change.
- **Lesson:** The principle that transfers beyond this platform.

## Decision map

**Identity and immutability**: The decisions that make a model change a routine, reversible operation instead of a project nobody can describe afterwards.

- ADR-01 · The embedding contract is the identity of an index, and an index is immutable under it
- ADR-02 · A contract change builds a new index beside the old one, with dual-write, and cuts over by alias
- ADR-03 · No query is ever answered from two contracts, because similarity scores are not comparable across models
- ADR-04 · The corpus stays with its owner; the chunk ledger is the platform's own system of record

**The cost of change**: The decisions that make the price of an edit proportional to what changed, and the price of a rebuild knowable in advance.

- ADR-05 · Chunk identity is a content hash of the normalised chunk text, not a position in the document
- ADR-06 · The chunker version is a contract element, so a chunking change is a full rebuild
- ADR-07 · A vector cache keyed by (chunk hash, contract), scoped to a tenant
- ADR-08 · Normalised text is retained for 30 days so a rebuild never touches the source systems

**Freshness and correctness**: The decisions about how fast the index follows the corpus, and what the platform refuses to show while it is catching up.

- ADR-09 · Three freshness lanes, with bulk on preemptible capacity and the batch wait published per lane
- ADR-10 · A document version becomes retrievable atomically, or not at all
- ADR-11 · The change log is keyed by document, at-least-once, with a monotonic source version as the arbiter
- ADR-12 · A reconciliation sweep is a first-class component, and the absence of a document never implies a deletion

**Permission and erasure**: The decisions about the most consequential fact in the system — who may see what — and about making a deletion true in every store.

- ADR-13 · Permissions are resolved at serve time against the owner's authority, and the platform fails closed
- ADR-14 · Revocation and deletion travel a read-path suppression list, not the indexing pipeline

**Quality and drift**: The decisions that let the platform say retrieval got worse, and say which of three indistinguishable causes the evidence supports.

- ADR-15 · Model replicas are addressed by digest, and every rollout is gated by a frozen reference probe
- ADR-16 · Cutover is gated on a frozen evaluation set, and drift is monitored as three separate causes

**Operating the fleet**: The decisions about running GPUs, degrading honestly, and recovering a store that is formally disposable and practically irreplaceable in the moment.

- ADR-17 · Indexes are snapshotted for recovery time, not durability, and a restore replays the deletion log

## Technology by capability

Open source, self-hosted on Kubernetes, chosen deliberately and for two reasons that point the same way. The first is rotation: the five most recent use cases in this practice all went to a hyperscaler, and a sixth would teach that cloud's service catalogue rather than architecture. The second is that this topic is one where the managed option hides the lesson. Behind a per-token embedding API, the cost of a model migration is a line item on an invoice; on an owned GPU fleet and an owned index it is a capacity plan, a rebuild window, a spot-versus-on-demand decision and a storage doubling the architecture has to absorb. Every requirement in ask.md is written vendor-neutrally — "durable change log", not a product name — so the table below is one defensible realisation rather than part of the requirement.

| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| Durable change log | Kafka, partitioned by document identifier, 13-month retention via tiered storage to MinIO | Open source, self-hosted | Redpanda; NATS JetStream; Pulsar; a managed stream service | Keyed partitioning gives per-document ordering, and long retention lets the transport double as the replay and deletion record | ADR-11 |
| Chunk ledger (system of record) | PostgreSQL with Patroni, synchronous standby in a second zone, partitioned by tenant | Open source, self-hosted | CockroachDB; YugabyteDB; a managed relational service | The one store with RPO 0, needing transactions, a cheap set-difference query for the diff, and operational familiarity | ADR-04 |
| Normalised text and snapshots | MinIO, erasure-coded, objects under tenant-scoped keys | Open source, self-hosted | Ceph RGW; SeaweedFS; cloud object storage | S3-compatible, cheap per TB, and the natural home for both retained text and index snapshots | ADR-08 |
| Vector cache | Redis, partitioned per tenant, keyed by (chunk hash, contract) | Open source, self-hosted | Dragonfly; Valkey; Memcached | A cost mechanism rather than a correctness one, so a simple, fast, losable store is exactly right | ADR-07 |
| Embedding inference | KubeRay with Ray Serve, running a text-embedding server, replicas pinned by image and weight digest | Open source, self-hosted | KServe; Triton; vLLM-based serving; a managed embedding API | Autoscaled replicas bound to node pools by taint, which is what makes three lanes with different capacity postures expressible | ADR-15 |
| Vector index | Qdrant, one collection per (tenant partition, contract), snapshots to MinIO | Open source, self-hosted | Milvus; Weaviate; pgvector; Vespa; a managed vector service | Cheap collection creation makes one-index-per-contract affordable, and payload filtering supports the version predicate atomic visibility needs | ADR-01 |
| Lexical index | OpenSearch over the same chunk rows | Open source, self-hosted | Elasticsearch; Tantivy; Vespa as a single hybrid engine | The degradation path when the vector tier is unavailable, and half of within-contract hybrid fusion | ADR-03 |
| Index aliases | etcd, compare-and-swap per (tenant, corpus), watched by the retrieval gateway | Open source, self-hosted | Consul; the relational catalogue itself; a vector store's native alias feature | A tiny, strongly consistent, watchable pointer store — exactly the shape of the alias, and separable from the catalogue's history | ADR-02 |
| Rebuild and migration orchestration | Argo Workflows and CronWorkflows, holding migration state in PostgreSQL | Open source, self-hosted | Temporal; Airflow; Dagster | Long-running, resumable, Kubernetes-native jobs for a 14-day rebuild, a nightly sweep and a restore | ADR-02 |
| Extraction and OCR | Tika-based parsers and Tesseract in egress-restricted pods, bounded in time, memory, depth and expansion ratio | Open source, self-hosted | Unstructured; a commercial document-intelligence API | The least-trusted compute in the system, so it must be isolated rather than merely sandboxed by convention | ADR-08 |
| Quality, drift and cost analytics | ClickHouse for evaluation runs, drift series and cost attribution | Open source, self-hosted | DuckDB on object storage; Druid; a cloud warehouse | Columnar scans over thirteen months of runs and series, which is a different access pattern from anything else in the platform | ADR-16 |
| Workload identity and secrets | SPIRE for SVIDs, OpenBao for source credentials, Keycloak for human identity | Open source, self-hosted | cert-manager with an internal CA; HashiCorp Vault; a cloud identity service | mTLS between every workload, short-lived source credentials rotatable without a pipeline outage, and operators separated from workloads | ADR-13 |
| Retrieval edge | Envoy, terminating mTLS and resolving the alias from a watched cache | Open source, self-hosted | NGINX; Traefik; a service mesh ingress | Resolving the alias at the edge from a cached pointer is what keeps a control-plane outage off the retrieval path | ADR-02 |
| Observability | OpenTelemetry collection, Prometheus with Mimir, Grafana, Loki and Tempo | Open source, self-hosted | VictoriaMetrics; a managed observability platform | Oldest-unembedded age per lane is both SLI and scaling trigger, so metrics have to be first-class rather than a side channel | ADR-09 |

## The decisions, and the alternatives that lost

### Identity and immutability

*The decisions that make a model change a routine, reversible operation instead of a project nobody can describe afterwards.*

#### ADR-01 · The embedding contract is the identity of an index, and an index is immutable under it

**Status:** Accepted  ·  **Shown on views:** 06, 12, 16, 19

*When a better embedding model ships, what exactly happens to the 4.8 billion vectors already in the index?*

**Context.** The obvious design treats the embedding model as configuration: a setting on the pipeline, pointing at a model name. It is simple, it needs no extra storage, and it has no answer to the question above. The honest answers available to it are all bad. Re-embed in place over weeks, and the index holds vectors from two models whose cosine similarities are being compared with each other — a correctness failure that produces no error and no alert, only slightly worse results. Re-embed in place over a weekend, and the cost is unbounded and the operation is unstoppable halfway. Or never re-embed, and the index is frozen on the model that happened to be current when the platform launched, while the field moves every few months. The same problem arrives from three other directions: a chunker improvement, a normaliser fix and a pooling change each invalidate stored vectors in exactly the same way, and a design that has no name for "the thing that produced this vector" cannot tell any of them apart.

**Decision.** An embedding contract is the tuple (normaliser version, chunker version, model digest, pooling and normalisation rule, dimensionality), and it is the identity of an index rather than configuration applied to one. Every stored vector records the contract that produced it; a vector's primary key is (chunk hash, contract id). The index builder refuses a vector whose contract does not match the index's. Changing any element of the contract therefore cannot modify an existing index — it can only produce a new one.

**How it is realised on AWS.** Contracts live in a PostgreSQL registry in the control plane, content-addressed by a hash of the tuple. The model digest is the registry digest of the served model image and weights, not a tag. Each tenant partition has one index per contract in Qdrant, named `<tenant>-<contract>-<n>`, and consumers never see that name: they resolve an alias held in etcd. The index builder reads the contract id from the catalogue when it opens a collection and rejects writes whose contract id differs, so the constraint is enforced at the only place vectors enter.

| Option | Verdict | Reasoning |
|---|---|---|
| Contract as index identity, indexes immutable under it | Chosen | Makes every contract element — model, chunker, normaliser, pooling — a first-class, priced, reversible change with one mechanism |
| Model as pipeline configuration, re-embedding in place | Rejected | Produces a heterogeneous index whose scores are meaningless across the heterogeneity, with no error and no way to describe the current state |
| Model version as a filterable field on the vector, one index for all contracts | Rejected | Keeps one index at the price of making every query specify a model and every ANN structure hold incomparable neighbours |
| Pin one model forever and never migrate | Right elsewhere | Right for a corpus that is loaded once and never improved on — a frozen archive, a regulatory snapshot. Wrong for a product whose search quality is the product |

**What it buys**

- A model, chunker or normaliser change is one operation with one shape: build beside, measure, flip the alias, retire the old index
- The state of the estate is always describable — every index names exactly what produced its contents, and every result can name the contract it came from
- Rollback exists by construction, because the previous contract's index is a separate artefact rather than overwritten rows
- Reproducibility is structural: given retained text and a contract id, any vector can be regenerated and compared

**What it costs**

- Vectors must be reproducible, so the chunk ledger and the retained normalised text become load-bearing stores with their own RPO rather than conveniences
- A full rebuild must fit a declared window and budget, which forces the fleet to be sized on the rebuild rate rather than the steady-state rate
- Two indexes exist during every migration, so the storage floor includes a planned doubling
- Nothing cheap is available for a small contract change: fixing a normaliser typo costs a full rebuild of the affected corpus

**Choose differently when.** If the corpus were static — loaded once, never edited, never re-embedded — the whole apparatus would be dead weight and a single mutable index would be correct. It would also flip if embedding models stopped improving, or if a model family emerged whose successive versions produced genuinely comparable vector spaces, which would make in-place upgrade safe. Neither looks likely inside this architecture's lifetime.

**Why it holds up over time.** The decision is stated in terms of identity and authority, not technology. It survives replacing Qdrant, replacing the GPU fleet with a managed embedding API, or moving to a vector representation that does not exist yet, because all of those are contract elements rather than architecture. What would date it is a change in the nature of embeddings themselves — multi-vector or late-interaction retrieval, where one chunk produces many vectors — and even then the contract concept extends rather than breaks.

> **Lesson.** When a derived store is produced by something that will change, give the producer a name and make that name part of the store's identity. A pipeline whose output cannot say what made it has no migration story, only a series of incidents.

#### ADR-02 · A contract change builds a new index beside the old one, with dual-write, and cuts over by alias

**Status:** Accepted  ·  **Shown on views:** 06, 16, 19

*How does a new index become the serving index without a window in which retrieval is wrong, missing or inconsistent?*

**Context.** ADR-01 establishes that a contract change produces a new index. That leaves the harder operational question of how it starts serving. The corpus does not stop being edited for the duration — at 8 million document versions a day, a 14-day rebuild sees over 100 million live edits land while it runs. An index built from a snapshot of the ledger and then switched to would be two weeks stale on the day it went live, and catching it up afterwards is the same problem again with less time. Meanwhile the old index must keep serving, correctly and freshly, throughout, because the migration is invisible to users and must stay that way.

**Decision.** A migration provisions the shadow index first, then enables dual-write, then begins re-embedding the ledger. From the moment dual-write is on, every live change is applied to both the serving and the shadow index under their own contracts. The shadow is therefore never stale on live edits, and "complete" means complete. Cutover is a write to the alias in the index catalogue. Rollback is the same write in reverse, and remains available for a declared 21-day retention period during which the superseded index continues to receive dual-write.

**How it is realised on AWS.** The migration orchestrator is an Argo Workflow holding the migration row in PostgreSQL. Enabling dual-write sets a flag the index-writer consumers read, after which each prepared chunk is embedded once per active contract and written to each matching index. Re-embedding walks the chunk ledger in tenant order through the bulk lane. Progress is published as the oldest unmigrated chunk age rather than a percentage, because a percentage hides a stalled tail. The alias flip is a single etcd compare-and-swap per tenant, and the retrieval gateway picks it up from a watch within a second.

| Option | Verdict | Reasoning |
|---|---|---|
| Shadow index with dual-write, atomic alias flip per tenant | Chosen | The shadow is current on live edits throughout, so cutover is instantaneous and rollback is symmetric |
| Build from a ledger snapshot, then catch up, then flip | Rejected | The catch-up is the same race with less margin, and the index is at its most stale exactly when it starts serving |
| Query-time fan-out across both contracts with score normalisation | Rejected | Requires comparing scores across models, which ADR-03 establishes is not honestly possible; also doubles query cost permanently |
| Blue/green at the cluster level rather than the index level | Right elsewhere | Reasonable where one tenant occupies a cluster. Here it would force all 25,000 organisations to flip together, removing progressive rollout |

**What it buys**

- Cutover is a single write, so there is no window in which retrieval is partially migrated
- Rollback is the same single write against an index that is still built, still fed and still measured — not an emergency path exercised once under pressure
- Progressive per-tenant flip becomes possible, which turns a 25,000-organisation migration into a series of small, observable ones
- The migration can be abandoned at any point by stopping dual-write and retiring the shadow, leaving no partial state

**What it costs**

- Peak storage during the overlap is roughly 1.6× steady state (60% additional, as the superseded index is retained 21 days)
- Every prepared chunk is embedded once per active contract while dual-write is on, which raises live-stream GPU cost for the window
- The orchestrator holds real state — which tenants are flipped, what dual-write is on — and becomes a component that must itself be correct
- A long migration means a long period in which the estate spans contracts, invalidating cross-tenant quality comparisons

**Choose differently when.** With a corpus small enough to rebuild inside a maintenance window — say under 50 million chunks at surge rate, a few hours — the whole dual-write apparatus is unnecessary and a stop-the-world rebuild is simpler and cheaper. The decision is justified by a rebuild measured in days against a corpus that is edited continuously.

**Why it holds up over time.** Build-beside-and-flip is one of the oldest patterns in operations and has outlived every technology it has been implemented on. What is specific here is dual-write during the overlap, and that follows from the corpus being live rather than from any component. It would only weaken if rebuild time fell to the point where staleness during a snapshot build stopped mattering.

> **Lesson.** A cutover you can reverse with the same action that performed it is a cutover you can afford to attempt. Build the reverse path first and the forward path becomes routine.

#### ADR-03 · No query is ever answered from two contracts, because similarity scores are not comparable across models

**Status:** Accepted  ·  **Shown on views:** 06, 15, 16

*During a migration, can a query search both the old and the new index and merge the results?*

**Context.** It is the first idea everyone has, and it looks like it solves the whole problem: if the new index is only 40% built, search both and merge, and the user gets the best available answer throughout. The difficulty is that the numbers being merged are not measurements of the same thing. Two embedding models place text in different spaces with different geometries; a cosine similarity of 0.82 in one is not stronger, weaker or equal to 0.82 in the other, and the relationship between them is not a monotonic function that can be calibrated away. Score normalisation across models — min-max over the candidate set, z-scoring, rank fusion — produces a number that looks comparable and is not, and the resulting ordering is an artefact of the normalisation rather than of relevance. Worse, the failure is silent and partial: most results are plausible, and the ones that are wrong look like ordinary retrieval noise.

**Decision.** A single query is resolved against exactly one contract. The alias a tenant resolves points at one index; there is no fan-out and no cross-contract merge. A tenant is either on the old contract or the new one, never both. Reciprocal-rank fusion is used within a contract to combine vector and lexical retrieval, where the two signals are deliberately different in kind and the fusion is over ranks rather than scores.

**How it is realised on AWS.** The alias in etcd maps (tenant, corpus) to exactly one index id. The retrieval gateway resolves it once per request and passes the contract id to the query embedder, so the query is embedded by the same model that produced the index. A contract mismatch between the query embedder and the index is a hard rejection rather than a degraded answer, and the rejection is counted.

| Option | Verdict | Reasoning |
|---|---|---|
| One contract per query, progressive per-tenant cutover | Chosen | Keeps every score comparison inside one space, and makes the migration unit a tenant rather than the estate |
| Fan-out across contracts with score normalisation | Rejected | Produces an ordering that is an artefact of the normalisation, failing silently and partially — the worst failure mode available |
| Fan-out with rank fusion instead of score normalisation | Rejected | Avoids the score problem and still merges two partial views of the corpus, so recall depends on migration progress in a way nobody can reason about |
| Fan-out during migration only, accepted as temporary | Right elsewhere | Defensible where retrieval is advisory and a user can see both result sets labelled by source. Not defensible behind an assistant that cites one answer |

**What it buys**

- Every score the platform compares lives in one space, so ranking means what it appears to mean
- Progressive per-tenant cutover becomes the natural migration shape, since the tenant is already the unit of index addressing
- A query embedder on the wrong contract is a loud rejection rather than a quiet quality regression
- Within-contract hybrid fusion stays legitimate, because rank fusion over two deliberately different signals is a different claim from score normalisation over two models

**What it costs**

- A tenant mid-migration sees results from the old contract until its flip, so the new model's benefit arrives in waves rather than gradually
- Cross-tenant comparisons of retrieval quality are invalid while the estate spans contracts, which complicates exactly the measurement a migration most wants
- The alias becomes per tenant rather than global, multiplying the number of pointers the catalogue holds

**Choose differently when.** If a model family shipped successive versions with provably aligned vector spaces — or if a cheap, reliable projection between two models' spaces existed and was itself evaluated — fan-out would become safe and the constraint would relax. Research in that direction exists; nothing production-ready does. It would also flip for a product where retrieval is advisory and the user can see which index a result came from.

**Why it holds up over time.** The underlying fact is mathematical rather than technological: two independently trained embedding spaces have no canonical alignment. That will not change because a vendor ships a new model. The decision should therefore outlast every component in this architecture, and its main risk is that someone re-proposes fan-out each time a migration feels slow.

> **Lesson.** Before merging two sets of numbers, ask whether they are measurements of the same thing. If they are not, a normalisation step does not fix it — it hides it, which is worse.

#### ADR-04 · The corpus stays with its owner; the chunk ledger is the platform's own system of record

**Status:** Accepted  ·  **Shown on views:** 01, 11, 12

*When the index and the corpus disagree about what a document says, which one is wrong — and what does the platform own well enough to rebuild from?*

**Context.** A retrieval platform sits in an awkward position: it must hold a derived representation of content it does not own, under permissions it does not control, and it must be able to recover from its own failures without asking 400 million documents to be re-delivered. Two tempting designs fail. Treating the index as authoritative makes a corrupt shard, a bad build or a wrong chunker unrecoverable except by full re-ingestion from source — a load spike the corpus owners will refuse. Copying the corpus wholesale into the platform makes recovery easy and creates a second system of record for content and permissions, with all the divergence, residency and erasure obligations that implies.

**Decision.** Content and permissions remain with their owners; the platform holds a pointer, a source version and a derived representation. The platform's own system of record is the chunk ledger: one row per (document, version, chunk hash) with offsets, structural path, contract version, vector reference and embedding status. Every vector index is a projection rebuildable from the chunk ledger, the retained normalised text and a pinned contract, with no input from any source system for documents inside the text retention window.

**How it is realised on AWS.** The chunk ledger is PostgreSQL with Patroni, synchronous standby in a second zone, RPO 0 and RTO 15 minutes — the only hard durability promise in the design. Normalised text sits in MinIO keyed by (document, version, normaliser version) with a 30-day expiry carried as a column rather than enforced by a cleanup job. A rebuild is an Argo Workflow that walks the ledger, reads text from MinIO, embeds through the bulk lane and writes a fresh index. Nothing but the ledger, the contract registry and the deletion log is backed up as truth.

| Option | Verdict | Reasoning |
|---|---|---|
| Corpus with its owner, chunk ledger as the platform's record, indexes as projections | Chosen | Recovers from any index failure without touching source systems, and keeps exactly one system of record for content |
| The index as the system of record | Rejected | Every index failure becomes a full re-ingestion from source — a load spike the corpus owners will not accept, and an RTO measured in days |
| Full corpus copy inside the platform | Rejected | Easy recovery at the price of a second source of truth for content and permissions, with the residency and erasure obligations that follow |
| No retained text; always re-fetch from source on rebuild | Right elsewhere | Right where the corpus is small or the source is cheap to read. At 400 million documents it makes every rebuild a source-system incident |

**What it buys**

- A corrupt shard, a bad build, a wrong chunker and a failed migration all have one remedy: rebuild from the ledger
- The platform's backup surface is small — a relational ledger, a registry and a log — rather than 12 TB of index data
- Erasure has a single authoritative place to start from, and reconciliation has something to compare against
- The derived representation can be thrown away deliberately, which is what makes ADR-01 affordable

**What it costs**

- The chunk ledger is a single-region relational store holding billions of rows and must be operated accordingly, with partitioning and vacuum pressure as real concerns
- Retained text is an additional store with its own retention policy, residency obligation and cost
- Outside the 30-day window, a rebuild does hit source systems, so the retention period is a bet about rebuild cadence
- Permissions resolved rather than held means a hot-path dependency on a system the platform does not own (ADR-13)

**Choose differently when.** If the corpus owner offered a durable, cheap, high-throughput read path with no load concern — or if the platform and the corpus were the same service — the retained-text store would be redundant and re-fetching on rebuild would be simpler. If the platform were instead the corpus owner, a single store would serve both purposes and this record would collapse into a schema decision.

**Why it holds up over time.** "The derived store is disposable and the ledger is not" is a statement about authority. It survives changing the ledger's technology, the text store's technology and the index's technology. The 30-day window is the part that will date, and it is stated as an assumption for that reason.

> **Lesson.** Own the smallest thing you cannot recreate, and make everything else recreatable from it. The size of what you must back up is a design output, not a fact of life.

### The cost of change

*The decisions that make the price of an edit proportional to what changed, and the price of a rebuild knowable in advance.*

#### ADR-05 · Chunk identity is a content hash of the normalised chunk text, not a position in the document

**Status:** Accepted  ·  **Shown on views:** 10, 12, 13

*When someone edits one paragraph of a 300-chunk document, how many chunks does the platform have to embed?*

**Context.** This single question decides whether the platform is affordable. At 8 million document versions a day, the difference between re-embedding every chunk of an edited document and re-embedding only the changed ones is roughly a factor of four in GPU spend and the difference between meeting and missing the interactive freshness budget. The naive identity — document plus ordinal — makes a paragraph inserted near the top shift every subsequent chunk's ordinal, so a one-word change invalidates the document. A byte-offset range has the same problem, shifted. The alternative, hashing the normalised chunk text, makes identity independent of position: unchanged text is the same chunk wherever it moved to. It brings its own complications, principally that the natural place to anchor a citation is a position rather than a hash, and that identical text appearing in thousands of documents collapses to one identity.

**Decision.** A chunk's identity is a hash of its normalised chunk text. Position — ordinal, offsets, structural path — is carried as attributes of the chunk row rather than as its key. The ledger diff between two document versions is therefore a set difference over hashes, and only hashes absent from the previous version enter the embedding stage. Chunk reuse rate is reported per corpus as a first-class operational metric.

**How it is realised on AWS.** The chunker emits, for each chunk, a SHA-256 over the normalised chunk text plus the chunker version. The ledger holds (chunk_hash PK, document_id, version, ordinal, offset_start, heading_path), so a chunk that moved keeps its hash and gets a new ordinal. The diff is a single query against the previous version's hash set. The vector cache (ADR-07) is keyed by the same hash, which is why a reverted edit costs nothing at all. Deduplication is scoped per tenant, so identical boilerplate collapses within an organisation but never across one.

| Option | Verdict | Reasoning |
|---|---|---|
| Content hash of the normalised chunk text | Chosen | Makes the cost of an edit proportional to what changed, and makes reverted edits and shared templates free |
| Document plus ordinal | Rejected | An insertion near the top invalidates every chunk below it, so a one-word edit costs a whole document |
| Document plus byte-offset range | Rejected | Same failure as the ordinal, and additionally fragile to any normaliser change that shifts offsets |
| Structural address (heading path plus index within section) | Right elsewhere | Attractive where citations must be permanently stable and documents are structurally rigid — a legal corpus, a standards library. Costs reuse on exactly the edits users make most |

**What it buys**

- A typical 41-chunk document edit produces about three changed chunks, making the edit stream affordable at 280 embeddings per second mean rather than roughly 1,100
- A reverted edit, a duplicated page and a shared template all resolve to cached vectors and cost no inference
- The diff is a cheap set operation rather than a text comparison, so preparation stays CPU-bound and predictable
- Reuse rate becomes a measurable, per-corpus number that predicts cost before a product ships on it

**What it costs**

- A citation cannot be anchored on the chunk key, so offsets must be carried separately and kept correct through normalisation — which is why ADR-08 retains the text that produced them
- Cross-document deduplication is deliberately limited to a tenant, giving up hit rate to avoid an existence side channel across an organisation boundary
- A chunk shared by many documents complicates deletion: the vector may only be dropped when the last referring document is gone
- A normaliser change alters every hash in the corpus, which is why the normaliser is a contract element and not a tweak

**Choose differently when.** For a corpus that is written once and never edited — ingested archives, published records — reuse is irrelevant and a positional identity would be simpler and would make citation anchoring trivial. The decision is justified entirely by continuous editing, and its value scales directly with edit rate.

**Why it holds up over time.** Content addressing is a durable idea and does not depend on the hash function, the store or the model. The chunker-version component of the hash is what keeps it honest as chunking evolves. What could date it is a move to multi-vector retrieval, where one chunk yields several vectors and identity has to extend to the vector within the chunk.

> **Lesson.** The identity you choose for a derived unit decides the cost of every future change to its source. Pick it by asking what a typical edit should cost, not by asking what is easiest to compute.

#### ADR-06 · The chunker version is a contract element, so a chunking change is a full rebuild

**Status:** Accepted  ·  **Shown on views:** 04, 12, 19

*Is improving the chunker a tuning change or a migration?*

**Context.** Chunking is the most-tinkered-with part of a retrieval system and the least respected. Teams change chunk size, overlap, boundary rules and context prefixes casually, because each change is a few lines and the effect on quality is immediately visible on a handful of queries. The consequence is almost never considered: a chunker change alters chunk boundaries, so it alters chunk text, so it alters every chunk hash, so every stored vector for that corpus is for text that no longer exists as a chunk. An index that has quietly accumulated three chunking regimes — depending on when each document was last edited — retrieves fragments of three different shapes and cannot be reasoned about. The temptation is to treat chunking as cheap because it needs no GPU; the reality is that a chunking change costs exactly as much as a model change, because both invalidate the same vectors.

**Decision.** The chunker version is an element of the embedding contract, with the same consequences as the model digest. A chunker change produces a new contract and therefore a new index, built beside and cut over by alias. Chunkers are selected per corpus type by the platform, not by the consuming team, and a chunker change is priced and gated exactly as a model change is.

**How it is realised on AWS.** The chunker is a versioned service; its version string enters the contract tuple and the chunk hash. Chunker configuration — size, overlap, boundary rules, context prefix composition — is part of the version rather than runtime configuration, so there is no path by which a deployed chunker behaves differently from the one that built the serving index. A chunker experiment runs by registering a candidate contract against one corpus and scoring it on that corpus's frozen evaluation set before any migration is proposed.

| Option | Verdict | Reasoning |
|---|---|---|
| Chunker version as a contract element | Chosen | Prices chunking honestly and keeps every stored vector reproducible from retained text |
| Chunker as runtime configuration | Rejected | Produces an index with several chunking regimes depending on last-edit date, and breaks reproducibility silently |
| Chunker version stored per chunk, mixed within one index | Rejected | Makes the heterogeneity visible without making it harmless: retrieval still compares fragments of different shapes |
| Consumer-chosen chunker per corpus | Right elsewhere | Right where consumers are sophisticated and corpora are few. Here it would hand fourteen teams a lever whose cost is a full rebuild they do not pay for |

**What it buys**

- Retrieval over any index is retrieval over chunks of one shape, which is a precondition for the scores meaning anything
- A chunker improvement is evaluated on a frozen query set before it is adopted, rather than shipped because it looked better on three examples
- Reproducibility holds: retained text plus a contract regenerates any vector exactly
- Chunking experiments become cheap at corpus scale and expensive at estate scale, which is the correct incentive

**What it costs**

- Chunking improvements are rare, because each is a 14-day background rebuild and a 21-day storage overlap for the affected corpus
- Consuming teams lose a lever they will expect to have, and the platform owes them an evaluation path instead
- The chunker becomes a release-managed component with its own versioning discipline, which is heavier than a config flag
- Normalisation and chunking are coupled through the hash, so a normaliser fix also triggers a rebuild

**Choose differently when.** If re-embedding were nearly free — a model small enough to run on CPU at corpus scale, or cached vectors that survive a chunking change, which they cannot — chunking could reasonably be runtime configuration. It also flips for a corpus small enough that a rebuild is minutes, where the discipline costs more than it saves.

**Why it holds up over time.** The coupling between chunk boundaries and stored vectors is inherent, not technological, so this holds as long as retrieval is over embedded chunks. It would be genuinely revisited by a retrieval architecture that embeds documents whole, or by late-interaction models where chunking is less determinative.

> **Lesson.** Find the cheap-looking change that invalidates expensive derived state, and give it the same ceremony as the expensive change. The accounting, not the diff size, decides what a change costs.

#### ADR-07 · A vector cache keyed by (chunk hash, contract), scoped to a tenant

**Status:** Accepted  ·  **Shown on views:** 10, 14, 18

*The same chunk text appears in a reverted edit, a copied page and a shared template. How many times is it embedded?*

**Context.** ADR-05 makes chunk identity content-derived, which immediately raises the possibility of never embedding the same text twice. The opportunity is substantial: in a collaboration corpus, template boilerplate, copied meeting notes, repeated headers and reverted edits mean a meaningful share of chunk hashes recur. The complication is scope. A cache shared across all tenants maximises hit rate and introduces an existence side channel — a caller who can observe a cache hit learns that some other organisation holds that exact text, which for a document fragment is a real disclosure. A cache scoped per tenant gives up cross-tenant hits and closes the channel.

**Decision.** A vector cache is maintained, keyed by (chunk content hash, contract id) and partitioned per tenant. A cache hit skips inference entirely. The cache is explicitly a cost mechanism rather than a correctness mechanism: it may be cold, partial or lost without affecting what the platform returns, only what it spends.

**How it is realised on AWS.** Redis, partitioned by tenant, holding the vector bytes against the composite key with a TTL long enough to cover an edit-and-revert cycle and a bulk re-run. The batcher consults the cache before forming a batch, so a fully cached document never reaches the GPU. Cache hit rate is a first-class operational signal alongside chunk reuse rate, because together they explain almost all variation in cost.

| Option | Verdict | Reasoning |
|---|---|---|
| Cache keyed by (chunk hash, contract), tenant-partitioned | Chosen | Captures the within-tenant recurrence that dominates, with no cross-tenant inference possible |
| Globally shared cache across all tenants | Rejected | Higher hit rate, and an observable existence side channel on exact document text across an organisation boundary |
| No cache; rely on the ledger diff alone | Rejected | Loses the free cases entirely — a reverted edit would cost full inference, which is the most common edit pattern after a save |
| Cache in the index rather than separately (read-before-write) | Right elsewhere | Sensible where the index is cheap to query by key. Here it puts build-path load on the serving tier, which is the one thing with a latency SLO |

**What it buys**

- A reverted edit costs nothing at all, which matters because save-then-undo is a normal editing pattern rather than an edge case
- Templates and duplicated pages within an organisation embed once, lowering the effective cost of large collaborative corpora
- Bulk re-runs after a transient failure are cheap, because the work already done is still cached
- Because the cache is a cost mechanism only, it can be resized, flushed or lost during an incident without any correctness review

**What it costs**

- Cross-tenant recurrence — which at 25,000 organisations is probably large — is deliberately given up
- Another stateful component to operate, size and monitor, holding derived data subject to the same tenant key controls as a vector anywhere else
- Hit rate becomes a cost dependency: a cache incident shows up as a GPU bill rather than as an outage, which is harder to notice
- TTL tuning is a real decision — too short loses the revert case, too long holds vectors for contracts being retired

**Choose differently when.** In a single-tenant deployment, or with a tenant population small and mutually trusting enough that content-existence disclosure is acceptable, a global cache would be correct and materially cheaper. The decision is driven by multi-tenancy, not by economics.

**Why it holds up over time.** Content-addressed caching of expensive derived values is stable. The tenant-scoping decision is a privacy judgement that should be revisited only with an explicit disclosure analysis, and the pressure to revisit it will grow as cost does — which is the main reason the reasoning is recorded here.

> **Lesson.** A deduplication key that spans a tenant boundary is a disclosure channel, however innocuous the value looks. Decide the scope of a cache on who could learn something from a hit, not only on what it saves.

#### ADR-08 · Normalised text is retained for 30 days so a rebuild never touches the source systems

**Status:** Accepted  ·  **Shown on views:** 10, 11, 12

*When the platform rebuilds an index, where does the text come from?*

**Context.** ADR-04 makes indexes rebuildable, and ADR-01 makes rebuilds routine rather than exceptional. That raises a question with an unattractive default answer: a rebuild that reads from source systems is a rebuild that hammers 400 million documents out of services that exist to serve users, plus a re-run of OCR on every scanned file. At 18,500 chunks a second surge rate, the source read load during a migration would dwarf normal product traffic. Corpus owners would quite reasonably refuse. Retaining the extracted text solves it, at the cost of holding a second copy of tenant content — with the residency, encryption and erasure obligations that implies, and a storage bill proportional to the corpus.

**Decision.** Normalised text is retained in object storage keyed by (document, version, normaliser version), for 30 days, with the expiry carried as data rather than enforced by a job. A rebuild reads retained text where it is inside the window and re-fetches from source only where it is not. The retained copy is treated with exactly the sensitivity of the source document, under tenant-scoped keys, and is erased on a deletion like any other store.

**How it is realised on AWS.** MinIO with erasure coding, objects encrypted under tenant-scoped keys, with `expires_at` on the corresponding ledger row so retention is visible and queryable rather than implicit in a lifecycle rule. A rebuild workflow partitions the ledger, reads text objects in bulk, and only falls back to the connector fetch path for documents whose text has expired — which it reports as a distinct count, because that count is the source-system load a migration will cause.

| Option | Verdict | Reasoning |
|---|---|---|
| Retain normalised text for 30 days; re-fetch only outside the window | Chosen | Makes re-chunking and re-embedding free of source load for recently touched documents, which is most of the working corpus |
| Retain nothing; always re-fetch and re-extract on rebuild | Rejected | Turns every migration into a source-system incident and re-runs OCR at corpus scale for no benefit |
| Retain indefinitely | Rejected | Removes the fallback path entirely, at the cost of a permanent full second copy of all tenant content and its obligations |
| Retain raw bytes rather than normalised text | Right elsewhere | Right if extraction itself is expected to change often. Here the normaliser is a contract element changed rarely, and raw bytes are larger and no more useful |

**What it buys**

- A chunker or model change re-enters the pipeline at the chunk stage, so a rebuild is GPU-bound and predictable rather than bounded by somebody else's API
- OCR is paid once per document version rather than once per rebuild, which for a scanned corpus is the dominant preparation cost
- The 30-day window covers the documents users actually touch, so in practice most of a rebuild reads from storage
- Expiry as data means retention is auditable and a privacy officer can query it rather than inferring it from a bucket policy

**What it costs**

- A second copy of tenant content exists, with residency, encryption, access-control and erasure obligations identical to the original
- Storage cost proportional to 30 days of document versions, which for an actively edited corpus is not small
- Rebuilds of cold documents still hit source systems, so the migration load profile depends on how old the corpus's tail is
- The window is a bet: a corpus whose rebuild cadence is longer than 30 days gets little benefit from it

**Choose differently when.** If extraction became nearly free and source systems offered a bulk, throttle-tolerant export path, retaining nothing would be cleaner and would remove a whole class of privacy obligation. Conversely, if migrations became frequent enough that the cold tail dominated, the window would need to grow — which is a cost decision, not an architectural one.

**Why it holds up over time.** The principle — keep the expensive intermediate, not the cheap one — is stable. The 30-day figure is explicitly an assumption and should be expected to move with observed rebuild cadence and with the cost ratio between storage and extraction.

> **Lesson.** When a derived store must be rebuildable, decide early which intermediate you keep. Rebuildability that depends on someone else's API is a promise you do not control.

### Freshness and correctness

*The decisions about how fast the index follows the corpus, and what the platform refuses to show while it is catching up.*

#### ADR-09 · Three freshness lanes, with bulk on preemptible capacity and the batch wait published per lane

**Status:** Accepted  ·  **Shown on views:** 05, 14, 17, 18

*GPU inference is far cheaper in large batches, and a large batch is bought with waiting. Who pays the wait?*

**Context.** The economics here are stark and frequently ignored. Embedding throughput per GPU rises steeply with batch size up to the point of saturation, so the difference between embedding chunks as they arrive and embedding them in large batches is a large multiple in cost per vector. Waiting is therefore the cheapest lever the platform has, and freshness is what it spends. Running one undifferentiated queue forces a single answer for everyone: either every edit waits, in which case the interactive promise is unmeetable, or nothing waits, in which case a 4.8-billion-chunk rebuild is unaffordable. Worse, one queue lets a single tenant's 400,000-document backfill sit in front of another tenant's edit from ten seconds ago.

**Decision.** Work is classified into three lanes at admission — interactive, standard and bulk — each with its own batch size, maximum wait and freshness commitment. The interactive lane has a short wait and accepts poor batch efficiency. The bulk lane has no freshness SLO, runs only on preemptible capacity, and is preempted rather than allowing interactive work to queue behind it. Each lane's maximum wait is a published number, so the freshness-for-cost trade is explicit rather than a tuning accident.

**How it is realised on AWS.** Admission assigns a lane from the corpus's declared freshness tier and the work's origin — a live change event is interactive or standard, a backfill or re-embedding run is always bulk. Each lane is a Kafka consumer group with its own batcher configuration. Ray Serve holds separate deployments bound to node pools by taint: interactive and standard on on-demand GPU nodes, bulk on spot nodes that spill to on-demand only when idle on-demand capacity exists. Oldest-unembedded age is tracked per lane and is both the freshness SLI and the autoscaling trigger.

| Option | Verdict | Reasoning |
|---|---|---|
| Three lanes with published per-lane batch wait, bulk on preemptible capacity | Chosen | Lets freshness be bought where it is perceived and cost be saved where it is not, with a 14-day rebuild made affordable by spot |
| One queue, batch for cost | Rejected | Makes the interactive promise unmeetable and lets one tenant's backfill delay another tenant's edit |
| One queue, no batching | Rejected | Meets freshness and makes a full rebuild cost several times its budget, which breaks the decision the whole architecture rests on |
| Per-tenant queues rather than per-lane | Right elsewhere | Stronger isolation, and at 25,000 tenants it fragments batches until batch efficiency collapses — right at hundreds of tenants, wrong at tens of thousands |

**What it buys**

- The interactive promise (p50 30 s) and the rebuild budget (14 days, ~$2,200 GPU) are simultaneously achievable, which no single-queue design manages
- A tenant's bulk backfill cannot delay another tenant's live edit, which is the isolation property that actually gets complained about
- Spot capacity withdrawal lengthens a migration and affects nothing else, so the cheapest capacity carries the most interruptible work
- Oldest-unembedded age per lane distinguishes a large backlog from a stuck one, which queue depth cannot

**What it costs**

- Three sets of batcher parameters and three node-pool postures to operate, tune and reason about
- Interactive batches are small and therefore expensive per vector; the architecture pays a premium precisely where users notice
- A migration's duration becomes dependent on spot availability, making completion estimates probabilistic
- Draining a large backlog deliberately breaches the standard lane to protect the interactive one — the right ordering, and still a breach

**Choose differently when.** If embedding inference became cheap enough that batch size stopped mattering — a much smaller model, or hardware where the batch curve flattens early — lanes would be unnecessary complexity and one queue would be correct. The whole structure exists to monetise waiting, and it disappears if waiting stops being worth anything.

**Why it holds up over time.** The batch-size-versus-latency curve is a property of accelerator hardware and has held across several generations. The specific rates will move by orders of magnitude; the shape of the trade, and therefore the need to assign the wait deliberately, should not.

> **Lesson.** When a resource is much cheaper in bulk, the question is not whether to batch but who pays the wait. Answer it per class of work, publish the answer, and the trade stops being an accident.

#### ADR-10 · A document version becomes retrievable atomically, or not at all

**Status:** Accepted  ·  **Shown on views:** 04, 10, 13

*While three of a document's chunks are being embedded, what does retrieval return for that document?*

**Context.** The convenient answer is: whatever is in the index. Chunks are independent rows, each upserted as its vector arrives, and retrieval simply sees a mixture — most chunks from version 40, the three changed ones from version 41. It costs nothing to implement and it produces the worst failure mode this platform has. A reader searching for text they just removed finds it, because the old chunk is still there. A reader searching for text they just added finds it, in a document whose surrounding chunks contradict it. The assistant cites two passages from the same document that disagree. None of this raises an error, and none of it is attributable: the document looks wrong in a way that reads as the model being bad. A stale document, by contrast, is a single coherent artefact that is simply behind — and the platform already has a mechanism for saying so.

**Decision.** A document version is retrievable only once every chunk belonging to that version is indexed. Until then, retrieval continues to serve the previous complete version. Visibility is a property recorded on the document version in the ledger, and the index excludes chunks of versions not yet marked visible.

**How it is realised on AWS.** The ledger's document_version row carries `visible_at`, set only when every chunk hash for that version has a vector in the target index. Chunks are written to the index tagged with their version, and retrieval filters on the visible version per document. Superseded chunks are removed in the same transition, so there is no window in which two versions are both visible. A version that cannot complete — a poison chunk, a repeated inference failure — never becomes visible and the document stays one version behind, which is reported as a distinct metric.

| Option | Verdict | Reasoning |
|---|---|---|
| Atomic per-document-version visibility | Chosen | A document is always one coherent version, and staleness is declarable rather than a silent mixture |
| Per-chunk visibility as vectors land | Rejected | Produces self-contradicting documents with no error, which reads to users as the retrieval being bad |
| Atomic across the whole corpus rather than per document | Rejected | Would make freshness for any document hostage to the slowest document in the corpus |
| Per-chunk visibility with a version filter applied by the consumer | Right elsewhere | Workable where there is one sophisticated consumer. With fourteen teams it moves a correctness invariant into fourteen codebases |

**What it buys**

- Retrieval over any document returns one internally consistent version, which is the precondition for a citation meaning anything
- Staleness becomes a single declarable fact per document — the version and its age — rather than an unknowable mixture
- A stuck document is visible as a document stuck at a version, which is a diagnosable condition rather than a quality complaint
- Deleted text actually stops matching at the moment the new version becomes visible

**What it costs**

- Freshness for a document is bounded by its slowest chunk, so a 310-chunk p99 document is slower than its median chunk suggests
- The index must support a version filter cheaply, which constrains the choice of vector store and adds a predicate to every query
- A poison chunk freezes a document at its previous version indefinitely until the quarantine is resolved
- The ledger becomes part of the visibility path, coupling retrieval correctness to a store that is otherwise build-path only

**Choose differently when.** If documents were small enough that a version was always one or two chunks, atomicity would be free and the decision trivial. It would also flip for a corpus of append-only records, where an older chunk can never contradict a newer one and a mixture is harmless.

**Why it holds up over time.** The underlying argument — a partially updated document is worse than a stale one — is about human interpretation, not technology, and does not weaken with scale. The implementation detail most likely to change is where the version filter is applied as vector stores grow richer predicate support.

> **Lesson.** Decide what the unit of consistency is before deciding how fast to update it. Speed applied to the wrong unit produces artefacts nobody can diagnose.

#### ADR-11 · The change log is keyed by document, at-least-once, with a monotonic source version as the arbiter

**Status:** Accepted  ·  **Shown on views:** 10, 13, 22

*Two edits to the same document arrive out of order. Which one ends up in the index?*

**Context.** Change feeds from real systems are at-least-once, occasionally out of order, and sometimes replayed in bulk after an incident at the source. The platform cannot ask for exactly-once delivery because no source offers it honestly. If order within a document is not guaranteed, the older edit can win and the index silently holds text the user replaced — the most expensive kind of wrong, because nobody notices until a search returns deleted wording. The platform also needs the ability to replay the log, both to recover from its own pipeline failures and to reprocess a corpus whose preparation was wrong, which means the log has to be a durable record rather than a transport.

**Decision.** Accepted changes are appended to a durable, replayable change log partitioned by document identifier, so all changes to one document are processed in order by one consumer. Every change carries a source-assigned monotonic version, and the pipeline refuses to apply a version lower than the one already recorded in the ledger. Delivery is assumed at-least-once; idempotence comes from the ledger's keys, not from the transport's promises.

**How it is realised on AWS.** Kafka with the document identifier as the partition key, 13-month retention through tiered storage to MinIO, which makes the log both the transport and the deletion-and-change record the compliance story depends on. The ledger holds the applied source version per document; a consumer that reads a lower version acknowledges and drops it, counting the drop. Replay is an Argo Workflow reading from an offset or from the archive.

| Option | Verdict | Reasoning |
|---|---|---|
| Document-keyed durable log, at-least-once, monotonic version arbiter | Chosen | Per-document ordering where it matters, idempotence in the data model, and replay for free |
| Trust the source feed's ordering and delivery | Rejected | No source offers this honestly, and the failure is a document silently holding replaced text |
| Timestamp rather than source version as the arbiter | Rejected | Clock skew between source and platform makes the arbiter non-monotonic exactly when two edits are close together |
| Global ordering across the whole corpus | Right elsewhere | Correct where cross-document transactions matter. Here it would serialise 8 million versions a day for a guarantee nothing needs |

**What it buys**

- Two rapid edits to one document cannot race, so the index converges on the newest version without a reconciliation pass
- A bulk replay from the source is harmless: every already-applied version is dropped idempotently
- The log doubles as the replayable record that makes a pipeline fix reprocessable and a deletion auditable
- No cross-document ordering is promised, which keeps throughput linear in partitions

**What it costs**

- A hot document — a page being collaboratively edited by twenty people — is a hot partition, and its throughput is one consumer's
- The platform depends on the source supplying a monotonic version; where one does not exist, the connector must synthesise one and the guarantee is only as good as that
- 13-month log retention is a storage commitment made for the compliance story rather than for the transport
- Nothing guarantees cross-document consistency, which occasionally surprises a consumer expecting a moved page and its parent to update together

**Choose differently when.** If every source offered a durable, ordered, replayable change feed of its own with a long retention, the platform's log would be a redundant copy and consuming the source feed directly would be simpler. In practice the log also serves the compliance and replay requirements, so it would have to be replaced by something rather than removed.

**Why it holds up over time.** At-least-once plus idempotence in the data model is the one delivery design that has survived every generation of messaging technology, precisely because it assumes the weakest guarantee. The monotonic-version arbiter is equally stable; what will change is the broker.

> **Lesson.** Do not buy ordering you do not need, and do not accept the absence of ordering you do. Partition by the unit whose internal order actually matters, and get idempotence from your own keys.

#### ADR-12 · A reconciliation sweep is a first-class component, and the absence of a document never implies a deletion

**Status:** Accepted  ·  **Shown on views:** 03, 10, 22

*A document was edited and no change event ever arrived. How does the platform find out?*

**Context.** Change feeds miss things. A webhook fails and is not retried, a connector is restarted mid-batch, a source deploys a bug, a tenant re-authorises an integration and the first hour is lost. Each of these is individually rare and collectively certain across 25,000 organisations and four source kinds. The failure is silent and permanent: the document sits in the index at an old version forever, and the only detector is a user noticing that search does not find something they wrote. The obvious fix — periodically compare the source's document inventory against the ledger — introduces its own hazard, because the same comparison that finds a missed edit also appears to find deletions. A source API that is partially unavailable, paginating badly or filtering by a permission the connector lost will under-report its own inventory, and a sweep that treats absence as deletion will then erase a tenant's corpus from the index.

**Decision.** A reconciliation sweep runs periodically per corpus, comparing the source's document inventory and versions against the chunk ledger. A document present in the source at a newer version than the ledger is re-ingested. A document absent from the source is never deleted on that basis: it is a candidate that must be confirmed by an explicit per-document check against the source, and a sweep whose inventory is implausibly smaller than the previous sweep's aborts rather than acting.

**How it is realised on AWS.** An Argo CronWorkflow per corpus, nightly by default and configurable per freshness tier. It pages the source inventory, diffs against the ledger by (document id, source version), and enqueues re-ingestion in the standard lane. Deletion candidates are checked individually through the connector's get-document path; a 404 confirms, anything else does not. A sweep aborts if the inventory count has fallen by more than a declared fraction since the last successful sweep, and the abort pages a human.

| Option | Verdict | Reasoning |
|---|---|---|
| Periodic sweep for missed edits, deletion only on explicit confirmation | Chosen | Catches feed gaps without giving a flaky source API the power to erase a tenant's retrievable corpus |
| No sweep; trust the feed | Rejected | Guarantees a slowly growing population of permanently stale documents, discovered only by users |
| Sweep that treats absence as deletion | Rejected | Converts any source-side pagination bug or permission change into mass removal from the index |
| Full periodic re-ingestion instead of a diff | Right elsewhere | Simpler and correct for a small corpus. At 400 million documents it is a continuous rebuild of everything |

**What it buys**

- A missed edit is corrected within one sweep interval rather than never, which bounds the worst case for freshness
- A flaky or partially-unavailable source cannot cause removal, because removal requires a positive confirmation
- Sweep discrepancy count becomes a direct measure of each source connector's reliability, which is otherwise invisible
- The abort threshold turns a catastrophic class of bug into a page rather than an incident

**What it costs**

- The sweep is real load on source inventory APIs, which has to be negotiated and rate-limited per corpus
- Up to one sweep interval of blindness remains, so the sweep is a bound on the failure rather than a fix for it
- Deletion confirmation is per document and therefore slow, so a genuine bulk deletion propagates through tombstones rather than through the sweep
- Another scheduled component whose own failure is silent unless its successful-run recency is itself monitored

**Choose differently when.** If a source offered a verifiable, gap-free change feed with a sequence number the platform could check for continuity, the sweep would be replaceable by a continuity check — far cheaper and strictly better. It would also change shape for a corpus small enough to re-inventory continuously.

**Why it holds up over time.** "Verify what you were told, and never infer destruction from silence" is a durable operational principle. The sweep's implementation will change with whatever the sources offer; the prohibition on treating absence as deletion should not.

> **Lesson.** Every event-driven integration needs a reconciliation path, and that path must be unable to cause the destructive action. Detection and deletion are different privileges.

### Permission and erasure

*The decisions about the most consequential fact in the system — who may see what — and about making a deletion true in every store.*

#### ADR-13 · Permissions are resolved at serve time against the owner's authority, and the platform fails closed

**Status:** Accepted  ·  **Shown on views:** 15, 20, 21

*Who decides whether this user may see this chunk, and what happens when that decision cannot be obtained?*

**Context.** This is the most consequential fact in the system and the one most often got wrong by copying. The tempting design stores an ACL alongside each vector and filters inside the index: one round trip, excellent recall, and no hot-path dependency. It fails on staleness. Sharing changes constantly in a collaboration product — a document moved between spaces, a group's membership edited, a guest removed — and an ACL copied at index time is wrong from the moment it is written. Worse, the copy is a reimplementation of someone else's permission model, including inheritance, group nesting, link sharing and per-space defaults, maintained by a team that does not own the semantics. Divergence is not a risk; it is the steady state. The alternative — ask the authority on every retrieval — is slower and introduces a dependency the platform cannot fail open on, because failing open means showing a user a document they were explicitly denied.

**Decision.** The platform stores an opaque reference to a document's permission state and resolves access against the corpus owner's authority at query time, in a batch, for the candidate set. Any failure of that resolution — timeout, error, ambiguity — excludes the affected chunks and the response reports that results were withheld. The platform never decides access itself and never caches an allow decision beyond the life of a request.

**How it is realised on AWS.** The retrieval gateway collects the distinct document identifiers in the candidate set, consults the read-path suppression list (ADR-14) first, then issues one batched authorisation call for the remainder, with a strict timeout. A partial response excludes the documents not positively allowed. The withheld count is returned to the consumer and recorded; the audit record carries counts, never content. Allow decisions are held only for the duration of the request.

| Option | Verdict | Reasoning |
|---|---|---|
| Serve-time batched resolution against the owner's authority, fail closed | Chosen | Always current and never a second source of truth for the fact that matters most |
| ACLs copied into the index, filtered pre-search | Rejected | Fast and stale: a reimplementation of someone else's permission semantics that diverges as its normal state |
| Physical index partitions per principal group | Rejected | Removes the hot-path dependency and generates an unbounded number of partitions under a real sharing model |
| Serve-time resolution with a short-lived decision cache | Right elsewhere | Right where revocation latency can be measured in minutes. Here the tenant admin's five-second promise forbids caching an allow |

**What it buys**

- Access is always evaluated against the current state of a model the platform does not own and does not have to track
- A revocation takes effect at the next query, with no index write and no propagation delay to reason about
- Fail-closed means the worst outcome of a dependency failure is a missing result, never a disclosure
- Withheld counts make the behaviour legible to a consumer, instead of results silently thinning

**What it costs**

- The authority is a hard hot-path dependency and therefore the real ceiling on retrieval availability, whatever the platform's own SLO says
- Post-filtering requires over-fetching candidates, and where a user may read a small fraction of a tenant's corpus, recall suffers measurably
- Batched authorisation adds latency to every query and makes retrieval p99 partly somebody else's number
- An authority incident presents as retrieval returning little or nothing, which is a confusing symptom for a product team to triage

**Choose differently when.** If the corpus owner published a change feed of permission events with a bounded propagation guarantee, a pushed permission cache would become defensible and would remove the hot-path dependency — this is the single most valuable thing a corpus owner could offer the platform. It also flips for a corpus where sharing is effectively static, such as a published handbook.

**Why it holds up over time.** Sharing models grow more complex over time, never less. A platform that copied a permission model in 2026 would be maintaining a divergent replica of an undocumented model by 2030. Asking is slower and stays correct, which is why this should outlast the components around it.

> **Lesson.** Never hold a second copy of the most consequential fact in your system. If the copy would be stale the moment you write it, resolve instead — and make the failure of resolution exclude rather than permit.

#### ADR-14 · Revocation and deletion travel a read-path suppression list, not the indexing pipeline

**Status:** Accepted  ·  **Shown on views:** 10, 20, 21

*A tenant admin revokes access to a document. How long until it stops being retrievable?*

**Context.** If revocation travels the same path as an edit, it inherits that path's latency: minutes at p50, ten minutes at p99, longer during a backlog. For an edit that is acceptable and honestly declarable. For a revocation it is not: the whole point of revoking access is that it takes effect now, and "the search index will catch up within ten minutes" is not an answer a privacy officer or an administrator will accept. The deeper problem is that erasure across every store — ledger, retained text, vector cache, several index versions, snapshots — is genuinely a distributed operation that cannot be instantaneous. So the platform needs two different promises: one about the read path, which can be fast, and one about erasure, which must be complete.

**Decision.** Deletion and permission revocation are written immediately to a suppression list consulted on the read path, before the permission batch, giving a 5-second p99 promise that the document stops being retrievable. Full erasure — chunk ledger, retained text, vector cache, every index version and any snapshot restored later — proceeds behind that, retried to completion, with tamper-evident evidence that it happened and an alert when a store has not confirmed inside the SLO.

**How it is realised on AWS.** Suppression is a small, replicated, read-optimised structure the retrieval gateway consults in-process with a short refresh interval; the write is synchronous on the admin path. Erasure is an Argo Workflow fanning out to each store, recording per-store confirmation in the deletion log. Any index snapshot restored later has the deletion log replayed over it before it serves traffic, which is what stops a restore resurrecting an erased document.

| Option | Verdict | Reasoning |
|---|---|---|
| Read-path suppression for immediacy, asynchronous erasure for completeness | Chosen | Two promises for two different obligations, neither compromised by the other |
| Revocation through the indexing pipeline | Rejected | Inherits the pipeline's latency and its backlog, which is the one thing a revocation must not do |
| Synchronous erasure across every store before acknowledging | Rejected | Makes the admin action as slow and as failure-prone as the slowest store, and offers no answer for a later snapshot restore |
| Rely on serve-time ACL resolution alone, with no suppression list | Right elsewhere | Defensible when the authority's own propagation is fast and a deletion is always visible to it. Here a deleted document may vanish from the authority too, and absence is not a denial |

**What it buys**

- Revocation is 5 seconds at p99 while indexing stays at minutes, and neither promise constrains the other
- Erasure is auditable per store rather than assumed, so "we deleted it" is evidence rather than a claim
- A snapshot restore cannot resurrect an erased document, which closes the most commonly missed erasure hole
- A store that fails to confirm raises an alert while the read path is already correct, so the incident is not a disclosure

**What it costs**

- A second mechanism on the read path, which must itself be highly available: suppression unavailable means fail closed, which is correct and still an outage
- The suppression list grows and needs its own compaction once erasure is confirmed across all stores
- Two promises to explain to consumers and auditors, and the distinction is easy to lose in a status page
- Bytes may remain present-but-unreachable while a retry is outstanding, which is correct on the read path and imperfect on disk

**Choose differently when.** If erasure across every store could be made genuinely fast — a far smaller corpus, or stores that all support cheap targeted deletion — one synchronous path would be simpler and the suppression list would be unnecessary complexity. The split exists because erasure at this scale cannot be fast and revocation must be.

**Why it holds up over time.** The separation of "stop showing it" from "remove every copy" is a permanent consequence of holding derived data in several stores, and regulatory expectations have moved consistently towards requiring both rather than one. This should hold.

> **Lesson.** When one action implies two obligations with incompatible time constants, build two mechanisms. Collapsing them means the fast promise is as slow as the slow one, or the slow one is quietly not kept.

### Quality and drift

*The decisions that let the platform say retrieval got worse, and say which of three indistinguishable causes the evidence supports.*

#### ADR-15 · Model replicas are addressed by digest, and every rollout is gated by a frozen reference probe

**Status:** Accepted  ·  **Shown on views:** 14, 19, 22

*How would the platform know that the model serving requests today is not the one that built the index?*

**Context.** This is the quietest catastrophic failure available to a retrieval system. A model is referenced by a name or a tag; a registry rebuild, a vendor update, a mutable tag pushed over, a different quantisation applied at serving time, or a kernel version that changes numerics slightly — and the vectors produced today are no longer in the same space as the vectors in the index. Nothing errors. Similarities drift by amounts too small to notice per query and large enough to degrade ranking across millions of them. The symptom arrives weeks later as "search feels worse", at which point the cause is indistinguishable from corpus drift, query drift, or a chunking change nobody logged. By then the index is an unknown mixture and the only remedy is a full rebuild.

**Decision.** Models are addressed by digest — of the image and the weights — never by a name or a tag, and the digest is an element of the embedding contract. Every model rollout embeds a frozen reference set of chunks and compares the output vectors against stored expected values. Divergence beyond a declared tolerance fails the rollout; it never fails the pipeline, which continues on the previous replicas.

**How it is realised on AWS.** The serving image and weights are pinned by content digest in the Ray Serve deployment, and the contract registry records that digest. A reference set of a few thousand chunks, with their expected vectors, is stored alongside the contract. On every replica start the probe runs before the replica is marked ready; a replica whose vectors diverge never receives traffic. The tolerance accounts for legitimate non-determinism in floating-point accumulation and nothing more.

| Option | Verdict | Reasoning |
|---|---|---|
| Digest addressing plus a frozen reference probe on every rollout | Chosen | Turns an invisible, delayed, unattributable failure into a failed deployment at the moment it is introduced |
| Model name or tag with a version label | Rejected | A mutable reference to the thing whose identity the entire index depends on |
| Digest addressing without a probe | Rejected | Catches a changed artefact and not a changed numeric environment — a different GPU, driver or kernel can move vectors with the same digest |
| Periodic probing rather than rollout gating | Right elsewhere | Better than nothing and detects the problem after some vectors are already wrong. Worth adding alongside, not instead |

**What it buys**

- A silently changed model becomes a failed rollout rather than a slow quality regression nobody can attribute
- The index's contents are genuinely identified by the contract, which is what makes ADR-01's guarantee real rather than nominal
- The probe also catches hardware and driver changes that move numerics without changing any artefact
- A rollout failure degrades nothing: the previous replicas keep serving, so the safe outcome is also the default one

**What it costs**

- Every rollout costs a probe before readiness, adding latency to deployments and a small amount of GPU time
- A tolerance has to be chosen, and choosing it badly produces either flapping rollouts or a probe that detects nothing
- The reference set must be representative; an unrepresentative set passes a model that moved in the part of the space that matters
- Model upgrades can no longer be done by moving a tag, which removes a convenience operators will miss

**Choose differently when.** Nothing plausible makes this unnecessary while the index depends on a model's exact numerics. If embeddings became robust to small numeric differences — or if retrieval moved to a representation where small shifts provably did not affect ranking — the probe could relax. Digest addressing would still be right.

**Why it holds up over time.** Content addressing of the artefact that produced persistent derived state is as durable a practice as exists in software. The probe is the part that adapts, since what needs checking depends on how much of the numeric environment the platform controls.

> **Lesson.** If a derived store depends on a computation's exact output, pin the computation by content and verify it on every deploy. A mutable reference to the thing your data depends on is a silent corruption waiting for a convenient moment.

#### ADR-16 · Cutover is gated on a frozen evaluation set, and drift is monitored as three separate causes

**Status:** Accepted  ·  **Shown on views:** 06, 18, 19

*Retrieval got worse. Did the corpus change, did the model change, or did the questions change?*

**Context.** All three produce the same complaint and the same dashboards, and only one of them is fixed by rolling back. A corpus that has drifted in subject matter — a company that moved into a new product line — makes old queries match worse because the relevant content is genuinely different. A model that changed makes everything match differently. A query mix that moved away from what the corpus covers makes retrieval look worse while the index is unchanged and correct. Treating all three as one metric means every alert triggers the same investigation and the same instinct, which is to roll back the last contract — correct a third of the time. The second difficulty is that none of this can rest on labelled relevance data, because nobody will fund labelling at corpus scale and a gate that needs labels will be skipped at the first deadline.

**Decision.** Each corpus maintains a frozen evaluation set of queries with recorded expected results, scored for recall@10 and MRR on a schedule and on every contract change. A contract cutover is gated on recall@10 not falling more than two percentage points against the outgoing contract; completeness alone never opens the gate. Separately, the distribution of stored vectors and the distribution of incoming queries are monitored independently, and every quality alert states which of the three causes the evidence supports, or that it cannot tell.

**How it is realised on AWS.** Evaluation sets are built from the real query log rather than from imagination, with judgements recorded once and reused across contracts. The quality harness runs in ClickHouse-backed batch jobs, scoring the set against a named index and storing the run against (contract, corpus). Vector distribution — norm, centroid, pairwise-similarity spread — and query distribution are tracked as separate time series with a 3σ alarm against a 7-day rolling baseline. Implicit signals (result selection, answer acceptance, reformulation, abandonment) are collected under a declared telemetry contract and attributed to a contract version.

| Option | Verdict | Reasoning |
|---|---|---|
| Frozen evaluation set as the gate, three drift signals monitored separately | Chosen | Gives a cutover an objective veto and makes a quality alert diagnosable rather than merely alarming |
| Gate on completeness; monitor quality afterwards | Rejected | Allows a fully built, measurably worse contract to start serving, with the regression found inside the rollback window only by luck |
| Gate on implicit signals alone | Rejected | Implicit signals lag by days and are confounded by product changes, so they cannot gate an operation that happens in seconds |
| Gate on human-labelled relevance judgements per migration | Right elsewhere | The strongest gate and the right one where a labelling budget exists. Here it would be skipped under deadline, which is worse than a weaker gate that is always run |

**What it buys**

- A migration can be refused on evidence, by a mechanism rather than by an argument
- Judgements are bought once and reused on every later contract, so the cost of the gate falls over time
- Separating the three drift causes means an alert points at an action instead of at a rollback reflex
- Thirteen months of evaluation history makes a slow regression across two consecutive migrations visible rather than lost at each cutover

**What it costs**

- The gate is exactly as good as the evaluation set, and a set drawn from an old query mix will certify a contract that is worse on current traffic
- Maintaining a frozen set per corpus is real work for fourteen consuming teams, and the platform owes them a harness to make it cheap
- Drift monitoring on 4.8 billion vectors is sampled, so a localised shift can be missed
- Implicit signals require a telemetry contract with consumers, which is a privacy and governance conversation as much as an engineering one

**Choose differently when.** If a reliable offline proxy for retrieval quality emerged — a model-based judge trustworthy enough to gate on — the frozen set would become a regression guard rather than the primary gate, and migrations would get faster. Conversely, if a corpus genuinely had no stable query distribution, the frozen set would stop being meaningful and the gate would have to be implicit-signal based with a slow, staged rollout.

**Why it holds up over time.** The three-causes distinction is a property of the problem and will not change. The gate mechanism should be expected to improve as evaluation methods do, and the architecture deliberately places the gate at a single point — the alias flip — so that improving it does not touch anything else.

> **Lesson.** When several different causes produce one symptom, instrument the causes separately before instrumenting the symptom. And prefer a weaker gate that is always run to a stronger one that is skipped under deadline.

### Operating the fleet

*The decisions about running GPUs, degrading honestly, and recovering a store that is formally disposable and practically irreplaceable in the moment.*

#### ADR-17 · Indexes are snapshotted for recovery time, not durability, and a restore replays the deletion log

**Status:** Accepted  ·  **Shown on views:** 11, 17, 22

*The vector index is formally disposable. Why back it up — and what makes a restore safe?*

**Context.** ADR-04 establishes that every index is a projection rebuildable from the ledger and retained text. The tempting conclusion is that indexes need no backup at all: if one is lost, rebuild it. The arithmetic refuses. A full rebuild runs 14 days at the background rate and 72 hours at surge, and no product owner will accept a 72-hour retrieval outage for a single-shard loss, let alone 14 days. So the rebuildability argument, which is correct about durability, says nothing useful about recovery time. Snapshots close that gap — and introduce a hazard of their own, because a snapshot taken before an erasure, restored after it, puts the erased document back into a serving index. That is a compliance failure produced by a recovery procedure, which is the kind that gets missed because the people who design restores and the people who design erasure are rarely the same people.

**Decision.** Vector indexes are snapshotted periodically to object storage, giving a 4-hour RTO, and the snapshots are explicitly a recovery-time mechanism rather than a durability claim — the ledger remains the record. Any restored snapshot has the deletion log replayed over it before it is allowed to serve traffic, and that replay is part of the restore procedure rather than a follow-up task.

**How it is realised on AWS.** Qdrant snapshots to MinIO on a schedule per collection, retained for a short window since their value decays. Restore is an Argo Workflow: fetch the snapshot, load it into a collection marked not-serving, replay the deletion log from the snapshot's timestamp forward, verify the suppression list against the collection, then make it serving by alias write. A restore that cannot complete the replay does not proceed; the fallback is a rebuild, slower and always correct.

| Option | Verdict | Reasoning |
|---|---|---|
| Snapshots for RTO, deletion-log replay before serving | Chosen | A 4-hour recovery that cannot resurrect erased content, with rebuild as the always-correct fallback |
| No snapshots; rebuild from the ledger on any loss | Rejected | Correct about durability and offers a 72-hour-to-14-day RTO, which no consumer will accept |
| Snapshots restored directly, with erasure reconciled afterwards | Rejected | Serves erased documents for the length of the reconciliation, which is a disclosure created by the recovery procedure |
| Continuous replication of the index instead of snapshots | Right elsewhere | Better RTO and right where the index is the record. Here it replicates a projection continuously and still needs the deletion replay on promotion |

**What it buys**

- Recovery time and durability are addressed by separate mechanisms, each honest about what it provides
- A restore cannot resurrect an erased document, because the replay is inside the procedure rather than after it
- Snapshot retention can be short, since the ledger is the record and an old snapshot has little value
- A failed replay degrades to a rebuild, so the unsafe path is never the fast path

**What it costs**

- Snapshot storage for 12 TB of index data, for a mechanism that exists only to shorten recovery
- The restore procedure is more complex than loading a snapshot and must be rehearsed, or it will be wrong when needed
- The deletion log becomes load-bearing for recovery as well as compliance, raising its own availability requirement
- A snapshot is a full copy of tenant-derived content and inherits every encryption, residency and access obligation

**Choose differently when.** If rebuild time fell inside an acceptable RTO — a much faster fleet, or vectors persisted in a columnar store so that a rebuild is IO-bound rather than GPU-bound, which is the Phase 3 option on the roadmap — snapshots would become unnecessary and this record would be superseded. That is the single change that would most simplify the platform's recovery story.

**Why it holds up over time.** The distinction between durability and recovery time is permanent and routinely conflated. The deletion-replay requirement will only strengthen as erasure obligations do. What may date is the need for snapshots at all, if rebuild becomes cheap.

> **Lesson.** "It is rebuildable" answers a durability question and says nothing about recovery time. Check which question your backup policy is actually answering — and make sure the restore path cannot undo a deletion.

## Every package used, in one table

The terms this package uses in a specific sense. Where a more common word exists and was rejected, the reason is given — usually because the common word implies something the architecture deliberately does not promise.

| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Embedding contract | The tuple (normaliser version, chunker version, model digest, pooling and normalisation rule, dimensionality). | The identity of an index and of every vector in it. The unit that is proposed, priced, built, gated, served and retired. | "Model version", which omits the chunker and normaliser and so understates what invalidates a vector. |
| Chunk hash | A hash over the normalised chunk text together with the chunker version. | Chunk identity, the key of the ledger diff, and the key of the vector cache. | "Chunk id", which suggests an assigned identifier and hides that identity is derived from content. |
| Chunk ledger | One relational row per (document, version, chunk hash) with offsets, structural path, contract version, vector reference and status. | The platform's own system of record, the basis of every reconciliation, and the input to every rebuild. | "Metadata store", which implies something incidental rather than the one store with RPO 0. |
| Alias | A mutable pointer from (tenant, corpus) to exactly one index id, with the previous id retained. | The only thing consumers resolve, and the mechanism of both cutover and rollback. | "Index name", which would couple every consumer to the contract currently in service. |
| Dual-write | The state in which every prepared chunk is embedded once per active contract and written to each matching index. | What keeps a shadow index current on live edits, so "complete" means complete and cutover is instantaneous. | "Backfill", which describes only the historical half and omits the live half that makes cutover safe. |
| Freshness lane | One of interactive, standard or bulk: a class of work with its own batch size, maximum wait, capacity posture and freshness commitment. | Where the batch-efficiency-for-freshness trade is made explicit and assigned. | "Priority", which implies ordering only and hides that each lane has a different capacity posture and a different promise. |
| Atomic version visibility | A document version is retrievable only once every chunk of that version is indexed. | What prevents a document matching on both its old and its new wording at once. | "Eventual consistency", which is true of the platform and says nothing about the unit over which consistency is guaranteed. |
| Suppression list | A read-path structure naming documents that must not be returned, consulted before the permission batch. | The 5-second revocation promise, held separate from the minutes-long indexing path and from the slower erasure fan-out. | "Deny list", which suggests a policy artefact rather than a time-critical correctness mechanism. |
| Reference probe | A frozen set of chunks with stored expected vectors, embedded on every model rollout and compared. | The detector for a model that changed behind an unchanged name, or a numeric environment that moved. | "Smoke test", which implies checking that the service answers rather than that it answers identically. |
| Drift | Three distinct conditions with one symptom: the corpus moved, the model moved, or the query distribution moved. | Monitored as three separate signals, because only one of the three is fixed by rolling back a contract. | Using "drift" for all three, which is how a quality alert becomes a rollback reflex that is right a third of the time. |
| Chunk reuse rate | The fraction of a document's chunks that are unchanged between two consecutive versions. | With cache hit rate, the number that sets the cost of the edit stream and sizes the embedding fleet. | "Cache hit rate" alone, which measures the second-order saving and misses the first-order one. |
| Rebuild | Reconstructing an index from the chunk ledger, the retained normalised text and a pinned contract, with no source-system input. | The single remedy for a corrupt index, a bad build, a wrong chunker and a failed migration. | "Restore", which implies returning to a previous state rather than deriving the current one. |
