# Architecture Decision Record

*Search Indexing Service · Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-10*

The argument these decisions serve is summarised in the [Architecture One-Pager](architecture-one-pager).

Sixteen decisions, grouped by the question they answer, each with the forcing question, the alternatives, what would flip the choice, and why it should outlast the technology it is realised on.

> **Status of this document.** This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a consumer marketplace with 40 million monthly active users, 180,000 active merchants and 12 product tenants, holding 80 million searchable documents across 40 indices and 220 GB of primary index data; 25,000 change events a second at steady state with a 4x ten-minute burst; 15,000 queries a second rising to 45,000 at the dinner-hour peak; and a largest index of 38 million documents that must rebuild inside four hours. Where the requirement supplied no number, one was invented and marked as an assumption in ask.md, section by section, so a reviewer can disagree with a figure and follow it to the decision that depends on it.

## How to read a record

- **Question:** The forcing question: why a decision was needed at all.
- **Context:** The requirement, the scale and the constraint that make it hard.
- **Decision:** What this architecture does, stated so it can be checked.
- **How it is realised on AWS:** The concrete mechanism: which service or package, configured how, in which subscription.
- **Options weighed:** Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- **Consequences:** What the choice buys and what it costs, both kept visible.
- **Choose differently when:** The conditions that would flip the decision for your system.
- **Why it holds up over time:** What keeps the decision right as scale, staff and technology change.
- **Lesson:** The principle that transfers beyond this platform.

## Decision map

**The alias boundary**: The decision that defines the architecture, and the two that follow from it directly.

- ADR-01 · Readers address an alias, never an index; every index is a disposable versioned projection
- ADR-02 · The change log is the platform's own system of record for what changed; sources remain the system of record for what is true
- ADR-03 · Breaking definition changes rebuild into a new index and swap; nothing is evolved in place

**Freshness**: Which staleness a user can catch, which they cannot, and who gets to see their own write.

- ADR-04 · Freshness is differentiated by field group into lanes with separate queues and separate budgets
- ADR-05 · Suppression fails closed with no tolerated staleness; enrichment and ranking signals fail open
- ADR-06 · The owner's read-your-writes is a read-path overlay, not a freshness requirement on the index

**Assembly**: How a document is built, what a shared-parent change costs, and why redelivery is harmless.

- ADR-07 · Documents are assembled and denormalised on write, and carry the source versions they were built from
- ADR-08 · A shared-parent change is a rate-limited, resumable, reportable fan-out job, not a side effect of one change
- ADR-09 · Idempotency lives in the data model: apply is guarded on (source, entity key, source version)

**Evidence before promotion**: What makes a new index or a new ranking good enough to serve, and how a lost document is found.

- ADR-10 · The swap gate is four independent machine-checkable detectors, and it refuses rather than escalating
- ADR-11 · Relevance is versioned configuration inside the index definition, split explicitly into index-time and query-time classes
- ADR-12 · A silently lost document is found by continuous sampled reconciliation, and a silent source by a quiet-period alarm

**Boundaries and placement**: What a query is allowed to see, who may move an alias, and where the platform runs.

- ADR-13 · Every volatile and access-controlling field is filtered before retrieval, from the caller's identity, never from the request
- ADR-14 · Moving an alias is a privilege granted separately from authoring a definition
- ADR-15 · One write region; read regions serve a follower index that is stale and says so
- ADR-16 · Shared cluster with per-tenant indices by default; dedicated clusters by declared policy, with live migration

## Technology by capability

Amazon Web Services was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases Microsoft Azure and self-hosted open source dominate, with AWS carrying a smaller share and Google Cloud smaller still — and reaching for the same cloud each time teaches a service catalogue rather than architecture. The second is fit: this topic genuinely wants a managed search cluster beside a managed durable log, and OpenSearch Service with MSK supplies both without the requirement becoming a product tour. Everything in ask.md is written vendor-neutrally — "durable change log", not a product name — and the table below is where the neutral capability meets a specific service, with what would be used instead on another stack.

| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| Change log — the platform's system of record for what changed | Amazon MSK, partitioned by hash of (source, entity key), 7 days hot with tiered archive to S3 for 90 days | Amazon Web Services | Confluent Cloud or self-managed Kafka; Azure Event Hubs with Capture; Google Pub/Sub with a replay subscription | Per-entity ordering without global ordering, retention long enough to cover the gap since the daily snapshot, and replay fast enough to rebuild the largest index inside the RTO | ADR-02 |
| Change capture from relational and key-value sources | AWS DMS for logical-decoding capture from Aurora PostgreSQL; DynamoDB Streams for the inventory source; API Gateway plus Lambda for the signed push path | Amazon Web Services | Debezium on Kubernetes; Azure Data Factory change feeds; Datastream on Google Cloud | Log-based capture rather than polling, because polling cannot see deletes and cannot be asked of a production database at this size | ADR-02 |
| Search indices, aliases and query execution | Amazon OpenSearch Service, one domain across three AZs, 220 GB primary with two replicas, sized to hold a second copy of the largest index | Amazon Web Services | Elastic Cloud; self-managed OpenSearch on EKS; Azure AI Search; Vespa for a learned-ranking future | Native atomic alias actions, partial document update, and per-index resource controls — the three engine features ADR-01, ADR-04 and ADR-16 all depend on | ADR-01 |
| Index definition registry, aliases, promotion history, judgement sets | Aurora PostgreSQL, multi-AZ with point-in-time recovery | Amazon Web Services | Cloud SQL or Spanner; Azure SQL; CockroachDB self-hosted | The smallest and most precious store in the package needs strong consistency, RPO 0 and real relational constraints over definition versions and alias state | ADR-11 |
| Assembly state — applied versions, join partners, parent mapping, enrichment cache | DynamoDB with conditional writes as the version guard, plus a 60-second-TTL recent-writes table for the owner overlay | Amazon Web Services | Bigtable; Azure Cosmos DB; Redis with persistence for the overlay only | A conditional write per (entity, source) is the version guard, and it must be enforced at the store rather than in process memory | ADR-09 |
| Assembly, writers, query service, reconciler | ECS Fargate services, scaled independently per lane, with reindex capacity rented for the duration of a rebuild | Amazon Web Services | EKS; Cloud Run; Azure Container Apps | Lane independence requires independent scaling, and a four-hour rebuild should rent capacity rather than own it | ADR-04 |
| Reindex orchestration with checkpoints and the acceptance gate | Step Functions driving snapshot replay, log catch-up, the four detectors and the alias action, with per-partition checkpoints in DynamoDB | Amazon Web Services | Temporal; Argo Workflows; Azure Durable Functions | A multi-hour, resumable, auditable process whose state must survive task replacement, and whose gate step must be able to refuse | ADR-10 |
| Snapshot store, archive, engagement log, quarantine, audit | S3 for daily source snapshots and retired-index archives; Kinesis Data Firehose to S3 with Athena for the query and engagement log; S3 Object Lock for the audit store | Amazon Web Services | GCS plus BigQuery; Azure Blob Storage with immutability policies plus Synapse | Immutability for audit, cheap retention for evidence, and query-in-place for relevance evaluation without standing infrastructure | ADR-14 |
| Identity, keys and the query-path perimeter | IAM roles per workload, KMS with tenant-scoped keys where obligations require, API Gateway for authn and quotas, WAF and CloudFront at the edge | Amazon Web Services | Workload identity on Kubernetes with Vault; Entra ID with Azure Front Door; Google Cloud Armor with Apigee | Four separable privileges — query, write a source, change a definition, move an alias — need four separable grants with per-workload identity | ADR-14 |
| Observability, freshness and cost signals | CloudWatch metrics and alarms with Managed Prometheus and Grafana for the end-to-end freshness and divergence dashboards; Cost and Usage Report tagged per tenant and index | Amazon Web Services | Self-hosted Prometheus, Mimir and Grafana; Azure Monitor; Google Cloud Monitoring | End-to-end per-lane freshness and a divergence rate are derived signals, not engine metrics, and the cost attribution has to reach per-index granularity | ADR-12 |

## The decisions, and the alternatives that lost

### The alias boundary

*The decision that defines the architecture, and the two that follow from it directly.*

#### ADR-01 · Readers address an alias, never an index; every index is a disposable versioned projection

**Status:** Accepted  ·  **Shown on views:** 02, 07, 14

*When the live index turns out to be wrong — wrong mapping, wrong analyser, corrupt shard, bad ranking — what is the recovery, and who has to change for it to work?*

**Context.** A search index accumulates reasons to be replaced faster than almost any other store: a field type was wrong, an analyser needs changing, a new scoring signal has to be materialised, a shard corrupted, a rebuild dropped documents, a relevance deploy made results worse. If readers name the index directly, each of those is a separate incident with a separate recovery, and several of them have no recovery at all — a mapping cannot be un-changed, and a document already written under the wrong analyser stays wrong until something rewrites it. The alternative is one indirection that costs nothing at read time: readers hold a name that the platform can repoint. Every search engine worth using offers some version of it, and the question is not whether the mechanism exists but whether the architecture is allowed to depend on it, because depending on it means no client code, no cached cursor and no runbook may ever contain a real index name.

**Decision.** All reads go through an alias. An index is an immutable-in-shape, versioned, disposable projection built from the change log and a source snapshot; making a new one live is an alias move, and rolling back is the same move in reverse. No reader-facing interface exposes an index name, and no operator procedure repairs a live index in place.

**How it is realised on AWS.** OpenSearch aliases are the reader-facing names, one per logical index per tenant, held in the Aurora definition registry as the authoritative record alongside the index_version row they point at and the promotion history. The query service resolves alias and relevance version per request from a short-TTL cache of the registry. Index names carry their definition version (idx_listings_v41), and the alias manager performs the swap as a single atomic alias action; the previous version stays mounted for 72 hours. Cursors issued for deep pagination are stamped with the index version that produced them and are rejected after a swap rather than silently re-run against a different index.

| Option | Verdict | Reasoning |
|---|---|---|
| Alias indirection with disposable versioned indices | Chosen | Collapses four unrelated failure modes into one routine operation with one rollback; costs double storage during a rebuild and a replay obligation on the log |
| Readers name the index directly | Rejected | Every schema change becomes a coordinated client release, and a bad mapping has no recovery that does not involve rewriting live data |
| A routing layer that rewrites index names per request | Rejected | Equivalent in effect but puts the indirection in a service the platform must keep up, rather than in the engine's own metadata where a swap is already atomic |
| Alias per variant, chosen by the client | Rejected | Leaks the experiment into every caller and makes a rollback a client change again; the variant split belongs behind the alias |

**What it buys**

- A breaking schema change, a wrong analyser, a corrupt shard and a bad relevance deploy are all handled by the same build-gate-swap path, with rollback as one alias move inside a minute
- Validation has somewhere to stand: a candidate index can be compared against the live one before any reader sees it
- Two relevance variants can serve behind one alias with a traffic split, so a ranking change is measurable on live traffic without a client release
- Capacity, mapping and analysis decisions stop being irreversible, which is what lets the platform accept a weekly rate of definition change

**What it costs**

- A second copy of the largest index during every rebuild — in this sizing, a planned capacity state rather than an exceptional one
- A replay obligation: the log and snapshots must be able to rebuild the largest index inside the RTO, which is a throughput commitment, not a backup policy
- Deep-pagination cursors are invalidated by a swap, so the query contract has to say so and clients have to handle it
- Index naming, alias state and promotion history become control-plane data that must be strongly consistent and backed up, even though the indices themselves are not

**Choose differently when.** If the document shape were genuinely fixed — a single-source index with no enrichment, no analysis changes and no ranking work, such as an internal lookup over a frozen reference dataset — the indirection would buy little and in-place updates would be the cheaper design. The decision is justified by a platform whose tenants expect to change a definition weekly and whose relevance is never finished.

**Why it holds up over time.** "Readers hold a pointer, never the thing" is a statement about where identity lives. An alias, a view, a routing rule or a symlink can each realise it, and a migration from one search engine to a successor is only possible at all because clients were never allowed to learn the index's real name. What would not survive is moving authority into the index itself — then there is nothing to rebuild from and no way back.

> **Lesson.** If a store is derived, make its identity private and its replacement routine. The cheapest reversibility in any architecture is one indirection installed before anyone needs it.

#### ADR-02 · The change log is the platform's own system of record for what changed; sources remain the system of record for what is true

**Status:** Accepted  ·  **Shown on views:** 10, 11, 14

*When an index has to be rebuilt, what is it rebuilt from — and what does that oblige the platform to retain?*

**Context.** A rebuild needs a base. The obvious one is the source systems themselves: scan the catalogue database, read every row, build the index. It is simple and it is also the thing source owners will refuse, because a full scan of a production database for 38 million rows is an outage risk they get nothing from, and because it cannot see what changed while the scan was running. The other obvious base is the index's own snapshots, which makes the index its own system of record and leaves no answer when the mapping inside the snapshot is the thing that was wrong. What is left is to retain the change stream, and to pair it with a periodic snapshot export that the source produces on its own schedule. That pairing is a platform commitment: the log has to reach back far enough to cover the gap since the snapshot, and replay fast enough to finish inside the recovery objective.

**Decision.** A durable, replicated, append-only change log partitioned by source and entity key is the platform's own system of record for what changed. A daily full snapshot in object storage is the rebuild base; the log covers the gap and the catch-up. The platform holds no authoritative copy of any business entity, and a search index is never a backup of anything.

**How it is realised on AWS.** Amazon MSK holds the change log, partitioned by a hash of (source, entity key) so per-entity ordering is preserved without global ordering, retained 7 days hot with tiered archive to S3 for 90 days. Sources export a daily full snapshot to S3 in a declared format; snapshot freshness per source is a published signal because a stale snapshot silently lengthens every rebuild. Reindex replays the snapshot through the same assembly code as the live path, then catches up from the log offset recorded with the snapshot. Log retention, snapshot cadence and replay throughput are tested together by a scheduled rehearsal rather than assumed.

| Option | Verdict | Reasoning |
|---|---|---|
| Retained change log plus periodic source snapshot | Chosen | Rebuild without touching production sources, and no gap between snapshot and stream; costs log retention and a replay-throughput commitment |
| Full scan of the sources on every rebuild | Rejected | Source owners carry all the risk for none of the benefit, and a scan cannot see concurrent changes without a second mechanism anyway |
| Index snapshots as the rebuild base | Rejected | Makes the index its own system of record, so a wrong mapping or a bad analyser is preserved faithfully into the restore |
| Log only, with no snapshot, retained from the beginning of time | Rejected | Elegant and unaffordable: replay time grows without bound and the rebuild budget is eventually consumed by history nobody disputes |

**What it buys**

- A rebuild never asks a source for anything it does not already produce, which is what makes rebuild-and-swap politically possible as well as technically possible
- A lost or wrong document can be repaired by replaying a key range rather than by a bespoke fix
- Snapshot plus log gives a defined, testable recovery path with a number attached, instead of a procedure nobody has run
- Because the log is the authority on what changed, late-arriving and reordered changes are a comparison rather than a correctness crisis

**What it costs**

- Log retention and archive are a standing cost that buys nothing on a good day
- Replay throughput becomes a capability to maintain and rehearse; it decays silently as documents get bigger and enrichment gets slower
- The platform depends on someone else's snapshot job, and a stale snapshot is a dependency failure with no symptom until a rebuild is attempted
- Two bases mean two ways to be wrong: a snapshot/log offset mismatch produces a gap that only reconciliation will find

**Choose differently when.** If every source could cheaply emit a consistent, fast full export on demand — a small dataset, a columnar store, a read replica nobody minds being hammered — the log could be retained for hours rather than days and the snapshot would do the work. The decision is justified by an 80-million-document estate whose sources are operational databases.

**Why it holds up over time.** The claim is that authority over "what changed" belongs to the platform and authority over "what is true" belongs to the source. Kafka, a managed successor, or an object-store-backed log can each hold it. What would not survive is the platform becoming the authority on an entity's current value, because then the index is no longer disposable and ADR-01 collapses with it.

> **Lesson.** A rebuildable store is only rebuildable if something retains the inputs. Decide what that something is, measure how fast it can replay, and rehearse it — otherwise "we can always rebuild" is a sentence rather than a capability.

#### ADR-03 · Breaking definition changes rebuild into a new index and swap; nothing is evolved in place

**Status:** Accepted  ·  **Shown on views:** 14, 18, 11

*A field's type is wrong, or an analyser has to change. Does the platform migrate the live index, or build a new one beside it?*

**Context.** Search engines allow some mapping evolution in place — adding a field is usually safe — and forbid the rest, which means a type change or an analysis change already requires rewriting every document. The tempting middle path is to do that rewrite inside the live index with versioned documents, migrating in the background: no double storage, no full replay. Its cost is that there is no moment at which the index is known-good, no clean place to run an acceptance check, a long tail of documents in two shapes being queried together, and no way back if the new shape is the problem. Building beside and swapping has the opposite profile: a second copy of the index for the duration, a full replay of snapshot plus log, and in exchange one atomic cutover, a free rollback, and somewhere to stand while validating.

**Decision.** Every definition change is classified as compatible or breaking at authoring time. A compatible change applies to the live index. A breaking change — a type change, an index-time analysis or synonym change, a removed or re-keyed field, a new materialised scoring signal — requires a full rebuild into a new index version, which becomes live only by passing the gate and moving the alias.

**How it is realised on AWS.** The definition API classifies the diff between two definition versions and refuses to apply a breaking change in place. A rebuild is a Step Functions execution: provision idx_vN+1 with the new mapping, replay the latest snapshot through assembly at a throttled write rate against a reserved capacity share, checkpoint per source partition in DynamoDB, catch up from the log offset, run the acceptance gate, then move the alias. The dry run reports projected duration, storage and cost from the document count and the measured replay rate before anyone commits to the change.

| Option | Verdict | Reasoning |
|---|---|---|
| Rebuild into a new index, gate, swap the alias | Chosen | Atomic cutover, free rollback, a validation point, and schema change as a routine operation; costs double storage for the window and a full replay |
| In-place mapping evolution with versioned documents | Rejected | Cheaper in storage and irreversible in effect: no known-good moment, no gate, and queries spanning two document shapes for the length of the migration |
| Dual-write to both indices during migration | Rejected | Halves the replay but doubles the write path and introduces a divergence class of its own; considered worth revisiting if replay throughput ever becomes the binding constraint |
| Rebuild only on a fixed schedule, batching definition changes | Rejected | Attractive for cost and terrible for ownership: it makes every definition change wait on the next window and couples unrelated tenants' changes into one risky swap |

**What it buys**

- Schema and analysis changes stop being one-way doors, which changes what tenants are willing to attempt
- A rebuild is the same operation whether the trigger is a definition change, a corruption, a rehearsal or a recovery — one code path, exercised constantly
- The dry run turns a definition change into a priced decision before it is made
- Because the same assembly code serves live writes and rebuilds, a rebuild validates the live path rather than exercising a parallel one

**What it costs**

- Cluster capacity must hold a second copy of the largest index, and the reserved reindex write share is capacity the query path cannot use
- A four-hour rebuild bounds how fast a breaking change can reach production, which is a real constraint on relevance work
- Rollback is only free for 72 hours; a regression found on day four needs a rebuild
- The classification itself becomes load-bearing: a change wrongly judged compatible is applied in place and is exactly the failure this decision exists to prevent

**Choose differently when.** If the index were small enough that a rebuild took minutes, the distinction would barely matter — rebuild everything, always, and skip the classification. If instead it were so large that a second copy were unaffordable, in-place evolution with dual-read would have to be reconsidered, and the rollback story replaced with something weaker and explicit. The decision sits in the middle band where a rebuild is hours and a second copy is affordable.

**Why it holds up over time.** "Replace, do not migrate" is a general property of derived stores, and it survives any particular engine's mapping rules. What would change it is an engine that can genuinely re-analyse documents in place with a transactional cutover — which would be the same decision realised differently, not a different decision.

> **Lesson.** Prefer replacing a derived store to migrating it. The second copy is usually cheaper than the incident, and it is certainly cheaper than having nowhere to validate.

### Freshness

*Which staleness a user can catch, which they cannot, and who gets to see their own write.*

#### ADR-04 · Freshness is differentiated by field group into lanes with separate queues and separate budgets

**Status:** Accepted  ·  **Shown on views:** 15, 10, 13

*Is there one freshness budget for a document, or several — and if several, what decides which field gets which?*

**Context.** One budget is simpler and is wrong in both directions at once. Set it at two seconds and the platform pays to re-assemble and re-enrich a whole document every time a stock count moves, which on a popular item happens hundreds of times an hour; the write cost is dominated by the fields that change most and matter least to assemble. Set it at a minute and a user gets told "closed" or "out of stock" one screen after a search that looked fine — the failure that makes people stop trusting a filter. The fields a user can instantly catch us about are a small, identifiable set, they are almost always the volatile ones, and they almost never need a join or an enrichment call. That coincidence is what makes two lanes affordable: the fast lane is cheap precisely because it does less work.

**Decision.** Each field group in a definition declares a freshness class. A fast lane carries the fields a user catches instantly — availability, stock, price, opening hours, moderation state — as partial document updates with no join and no enrichment. A slow lane carries titles, descriptions, attributes and computed signals as whole-document writes. The lanes have separate queues, separate consumer groups and separate budgets, and neither can consume the other's capacity.

**How it is realised on AWS.** Separate MSK topics per lane, fed by the capture tier from the field group a change touches, with separate ECS assembly and writer services and separate scaling policies. Fast-lane writes are field-level partial updates with a version guard per field; slow-lane writes replace the whole document. Freshness is measured end to end per lane — source commit to searchable — rather than as consumer lag, because consumer lag can read zero while the index is wrong. Both lanes' observed freshness is published and stamped on query responses.

| Option | Verdict | Reasoning |
|---|---|---|
| Two lanes plus a suppression lane, partial update on the fast lane | Chosen | Spends freshness where it is perceived and write cost where it is cheap; costs two consistency stories inside one document |
| One lane, one budget, whole-document writes | Rejected | Either unaffordable or visibly stale; the write cost is set by the most volatile field and the perceived quality by the least fresh one |
| Volatile fields not indexed at all, joined at query time from a live store | Rejected | Keeps the index stable and makes every query depend on another store's availability and latency; it also cannot filter before retrieval, which ADR-13 forbids |
| Per-tenant single budget, tenant chooses the number | Rejected | Pushes a cost decision onto someone with no visibility of it, and produces tenants who all choose two seconds |

**What it buys**

- A 4x burst of slow-lane edits, or a four-hour rebuild, cannot move the fast lane's p99
- Write cost falls sharply, because the highest-frequency changes skip joins and enrichment entirely
- The budget per class is a published number a product team can design a UI around
- The classification is a definition-time decision, which makes it reviewable rather than emergent

**What it costs**

- Two lanes writing the same document create a collision class: a partial update can land on a document the slow lane is about to replace, so field-level version guards are mandatory rather than an optimisation
- A reader holding one document now holds fields with two different staleness properties, and the response can only report the index's freshness, not the field's
- More moving parts: two pipelines to scale, monitor and reason about, and a field group whose class is wrong is invisible until a user notices
- Reconciliation must assert specifically on fast-lane fields, because that is where a silent revert would hide

**Choose differently when.** If the volatile fields were not also the cheap ones — if availability required a three-source join and a model call — the fast lane would stop being cheap and the right answer would move towards a query-time lookup. The decision rests on an observed property of this domain: what users catch is what is simple to write.

**Why it holds up over time.** The ranking of fields by human perceptibility does not change when hardware does. Faster clusters tighten both budgets and leave the split intact. What would end the decision is a store that could offer per-field freshness guarantees natively, which would be this decision implemented by someone else.

> **Lesson.** Freshness is not a property of a system; it is a property of a field, and the budget should follow what a user can catch rather than what is convenient to average.

#### ADR-05 · Suppression fails closed with no tolerated staleness; enrichment and ranking signals fail open

**Status:** Accepted  ·  **Shown on views:** 15, 19, 21

*When part of the pipeline is broken or behind, which kind of wrongness is the platform allowed to serve?*

**Context.** Every indexing platform is behind on something at any moment, so the useful question is not whether to be wrong but which wrongness to prefer. Two classes behave very differently. Showing something that should no longer exist — a deleted listing, a suspended merchant, a moderated item, something the caller is not permitted to see — is a wrongness with no acceptable size: there is no filter under which the user can have it, and in the access case it is a disclosure. Showing something with a stale category, a stale popularity score or a stale geocode is a wrongness the user cannot usually detect and that costs some relevance. Treating both the same produces either a platform that withholds documents because a scoring service is down, or one that serves suspended listings because a moderation event is queued behind a price storm.

**Decision.** Suppression, deletion, moderation state and access constraints fail closed: they have their own priority lane, no tolerated staleness beyond a 5-second p99, and a document is removed from the candidate set rather than ranked down. Enrichment values and ranking signals fail open: on timeout or error the document is indexed with its previous value and flagged degraded, and is never withheld from the index for want of an enrichment.

**How it is realised on AWS.** A suppression topic and writer ahead of the fast lane, so a moderation action cannot queue behind a price storm. Suppression and access fields are indexed and applied as mandatory pre-retrieval filters (ADR-13). Enrichment calls carry a per-service timeout and a previous-value fallback read from the DynamoDB enrichment cache; the document records a degraded flag per enrichment, which is published as a share of indexed documents and backfilled when the service returns.

| Option | Verdict | Reasoning |
|---|---|---|
| Fail closed on suppression, open on enrichment | Chosen | Matches the asymmetry of the harm; costs a third lane and a degraded-document state that has to be visible |
| Fail closed on everything | Rejected | An enrichment outage becomes a search outage: documents withheld because a categoriser is down, which no user would choose |
| Fail open on everything | Rejected | Suspended and deleted items stay searchable while a queue drains, and access filters degrade into a disclosure |
| Rank suppressed items down instead of removing them | Rejected | Keeps them in facet counts and reachable by pagination or a narrow filter, which is the same failure with more steps |

**What it buys**

- The one staleness class users cannot tolerate is the one with the tightest budget and its own capacity
- An enrichment dependency outage is a relevance incident rather than an availability incident
- "Degraded" is a visible, measurable state, so quality loss is observable rather than inferred
- The rule is simple enough to apply to a new field group at definition time without a meeting

**What it costs**

- A third lane to operate, scale and monitor, whose normal state is almost empty — which makes its own failure easy to miss
- Documents can sit indefinitely on a stale enrichment value if a service stays down, decaying relevance invisibly unless the degraded share is watched
- Fail-closed on access means an identity-service problem can make legitimate documents unreachable, which is the correct trade and still an outage
- Backfill after an enrichment outage is a fan-out-shaped workload that competes with the slow lane

**Choose differently when.** If the platform served only public, non-moderated, non-perishable content — a documentation search, a public reference corpus — the asymmetry would disappear and one lane would do. The decision is justified by an index over commerce and user-generated content, where suppression has legal and safety weight.

**Why it holds up over time.** The asymmetry between "showed something that should be gone" and "ranked something slightly wrong" is about consequences, not implementation, and does not change with the stack. What would change is the list of fields in each class, which is why the class is declared per field group rather than hard-coded.

> **Lesson.** Decide the direction of failure per field, not per system. Absence is usually recoverable; a wrong presence often is not.

#### ADR-06 · The owner's read-your-writes is a read-path overlay, not a freshness requirement on the index

**Status:** Accepted  ·  **Shown on views:** 05, 20, 13

*A merchant saves a price and immediately searches their own shop. Who pays for them seeing the new value?*

**Context.** This is the single most common complaint about any indexing platform, and the most expensive one to answer naively. The person who just made a change checks immediately, knows exactly what they changed, and treats any delay as the system losing their work. One merchant's expectation, taken literally, becomes a sub-second freshness requirement on the whole index — which would force synchronous indexing on the write path and discard the economics of the entire design. But the requirement is narrower than it looks: it applies to one principal, their own entities, and a window of seconds. It does not apply to the eighty million documents everyone else is searching.

**Decision.** Read-your-writes is guaranteed for the entity's owner on their own query within one second, delivered as an overlay in the query service reading a short-TTL recent-writes table. It is not a general consistency mechanism, is not offered on consumer queries, and places no additional freshness obligation on the index itself.

**How it is realised on AWS.** On accepting a change, the capture tier writes a compact record — principal, entity key, changed fields, source version — to a DynamoDB table with a 60-second TTL. A query carrying an authenticated principal has its results post-processed against that table for entities the principal owns: a changed field is overlaid, and an entity that would now match but does not yet appear is injected when the query's filters allow it. The overlay never applies across principals, never alters facet counts, and declares itself in the response alongside the index version and observed freshness.

| Option | Verdict | Reasoning |
|---|---|---|
| Read-path overlay from a short-TTL recent-writes table | Chosen | Satisfies the one principal who notices without changing the index's freshness budget; costs an extra read on authenticated queries and an overlay that must not leak across principals |
| Synchronous indexing on the owner's write path | Rejected | Puts the search cluster on the critical path of every merchant save, and makes a cluster problem a write outage |
| Session-pinned version token the query waits for | Rejected | Honest and slow: it converts the pipeline's tail latency directly into the merchant's query latency, and needs client state the console may not carry |
| Tell the owner it takes up to a minute | Rejected | Cheapest and genuinely viable for some products; rejected here because the merchant console's whole job is to make a change feel applied |

**What it buys**

- The expensive expectation is met for the only person who holds it, at the cost of one key-value read on authenticated queries
- The index's freshness budget stays a platform decision rather than being set by the most impatient user
- The overlay doubles as an honest answer to "did my change land": the record exists the moment the change is durable
- Because it is read-path, it works identically during a rebuild, a lane backlog or a regional failover

**What it costs**

- An overlay is a second code path for producing a result, and a bug in it is a cross-principal data leak rather than a latency problem
- Facet counts and pagination are not overlaid, so an owner can see their changed item in a result list whose counts disagree with it by one
- A 60-second TTL is a guess: a lane backlog longer than the TTL reopens the complaint exactly when the platform is already behind
- It is only available where the caller is authenticated and owns the entity, which excludes the consumer case entirely — correctly, and visibly

**Choose differently when.** If the fast lane could credibly guarantee sub-second end-to-end freshness for every document, the overlay would be unnecessary complexity. If instead the product had no authenticated authoring surface at all, the requirement would not exist. The decision lives in the gap between a 2-second index budget and a 1-second human expectation.

**Why it holds up over time.** "Spend consistency where it is perceived" is a design posture, not a technology. The overlay could be realised in a cache, a sidecar, an engine feature or a client-side merge; what matters is that the obligation stays attached to one principal and does not propagate into the index's budget.

> **Lesson.** When one user's expectation would set a system-wide requirement, check whether it is actually a requirement about one user. It usually is, and it is usually cheaper to serve it on the read path.

### Assembly

*How a document is built, what a shared-parent change costs, and why redelivery is harmless.*

#### ADR-07 · Documents are assembled and denormalised on write, and carry the source versions they were built from

**Status:** Accepted  ·  **Shown on views:** 10, 12, 15

*Does the platform build one flat document per searchable thing, or keep the pieces separate and join them when a query arrives?*

**Context.** Search engines reward flat documents: one document, one scoring pass, filters and facets over fields that are physically present. The cost is that a change to anything a document contains means rewriting the document, and a change to something many documents contain means rewriting many. The alternative — parent-child or nested mappings, or a runtime lookup against a small mutable store — makes the shared change cheap and spends the difference on every query, in latency and in query complexity, with filters over looked-up fields becoming either impossible or post-retrieval. Because this platform's reads outnumber its writes and its latency budget is tight, and because ADR-13 requires every volatile field to be filterable before retrieval, the flat document is the shape that satisfies the constraints that cannot be relaxed.

**Decision.** Assembly produces whole, flat, self-describing documents on write: sources joined on declared keys, enrichment applied, and the document stamped with its assembly version, the source versions it was built from, and any degradation flags. Query-time joins are not used for anything that must be filtered, faceted or sorted on.

**How it is realised on AWS.** ECS assembly workers consume the change log, read current source versions and join partners from the DynamoDB assembly-state table, call enrichment with timeouts, and emit a complete document. Assembly is deterministic — the same source versions and definition version produce a byte-identical document — which is what makes a candidate index diffable against the live one in the swap gate. A change arriving for an entity whose join partner has not yet been seen is indexed in a declared incomplete state naming the missing sources, and re-emitted when the partner arrives.

| Option | Verdict | Reasoning |
|---|---|---|
| Denormalise on write into flat self-describing documents | Chosen | Trivial query, filterable volatile fields, diffable rebuilds; costs write amplification and a fan-out problem that ADR-08 has to solve |
| Parent-child or nested mapping, join inside the engine | Rejected | Removes the fan-out and pays on every query; nested filters and facets are materially slower and more fragile at this query rate |
| Runtime lookup of volatile fields from a key-value store | Rejected | Cannot filter before retrieval, so pagination and facet counts break — the failure view 4 identifies as the one users notice |
| Hybrid: flat for stable fields, looked up for volatile ones | Rejected | A real option with a real cost — two consistency models in one response and a second store on the query path; superseded here by the fast lane, which keeps volatile fields in the index cheaply |

**What it buys**

- One scoring pass over physically present fields keeps the p99 inside 250 ms at 15,000 queries a second
- Every volatile field is filterable before retrieval, which is what makes ADR-13 implementable at all
- Deterministic assembly makes a rebuild comparable to the live index, which is the only reason the sampled-diff detector in the gate can exist
- A document that records its inputs can be reconciled against them, which is the basis of ADR-12

**What it costs**

- Write amplification: a change to a shared parent rewrites every child, which is the whole of ADR-08
- Document size grows with denormalisation, and so does index storage and bulk-write cost
- An incomplete document is a real state that queries can see, and the "missing source" flag has to mean something to a caller
- Determinism is a constraint on assembly code forever: any non-deterministic enrichment or ordering silently disables the diff detector

**Choose differently when.** If the read-to-write ratio inverted — an index written constantly and queried rarely, such as an operational audit search — the fan-out cost would dominate and a nested or lookup model would win. The decision is justified by roughly 15,000 queries a second against 25,000 changes a second spread over 80 million documents, where any one document is read far more often than it is written.

**Why it holds up over time.** "Pay once on write for a read you will do many times" is an economic argument that survives engines. What would change it is an engine where joins are genuinely free at query time, or a workload where writes outnumber reads — either of which would be a different decision made on the same grounds.

> **Lesson.** Decide the document shape from the reads you cannot relax, not from the writes you would like to make cheap. Then budget separately for the write pattern that shape creates.

#### ADR-08 · A shared-parent change is a rate-limited, resumable, reportable fan-out job, not a side effect of one change

**Status:** Accepted  ·  **Shown on views:** 15, 12, 17

*Catalogue ops renames a brand that appears on two million listings. What happens in the next five minutes?*

**Context.** Denormalisation makes this change expensive by construction (ADR-07). Treated as an ordinary change, it enters the pipeline as one event and leaves as two million writes, which saturates the assembly tier, floods the cluster's bulk queue, blows every lane's freshness budget and — because nothing is tracking it — gives nobody an answer to "how long will this take?". It also cannot be made atomic at any affordable price, so the honest question is not how to avoid a partial state but which partial state to serve and how to describe it. Mid-fan-out, a user searching will see some documents with the old brand and some with the new, and that is either an accepted, bounded, reported condition or an unexplained inconsistency.

**Decision.** A change to a shared parent entity is resolved into the set of affected children and executed as a tracked background job: rate-limited, checkpointed, resumable, with progress and an ETA. A single parent change may enqueue no more than a declared number of child rebuilds synchronously; above that threshold it becomes a job. Mid-fan-out inconsistency is accepted and bounded by a published completion target rather than hidden.

**How it is realised on AWS.** The assembly-state table holds the parent-to-children mapping, so the child set is a lookup rather than a scan. Above the synchronous threshold, a Step Functions execution drives batched child rebuilds through SQS at a configured rate against the slow lane's capacity, checkpointing progress; the job's children_total, children_done and rate limit are a tenant-visible record, and fan-out backlog is a first-class signal and an autoscaling input. Jobs are preemptible by the fast and suppression lanes and resume after a failure from their checkpoint.

| Option | Verdict | Reasoning |
|---|---|---|
| Rate-limited resumable tracked job with bounded synchronous enqueue | Chosen | Bounds the blast radius and gives an answer to "how long"; costs an orchestration path and an accepted window of inconsistency |
| Eager fan-out through the normal pipeline | Rejected | One edit becomes a self-inflicted 4x burst with no progress reporting and no way to slow it down |
| Parent-child mapping so the parent is stored once | Rejected | Genuinely removes the problem and reintroduces the query-time join ADR-07 rejected; the strongest alternative here |
| Defer parent changes to the next scheduled rebuild | Rejected | Cheap and unacceptable: a brand rename invisible for hours is the kind of staleness catalogue ops is explicitly trying to avoid |

**What it buys**

- A two-million-document edit has a completion target (30 minutes at p95) and a progress surface instead of a rumour
- The fast and suppression lanes keep their budgets during a bulk edit, because the job runs against the slow lane's capacity and is preemptible
- Resumability means a failed worker costs a checkpoint interval, not a restart of the whole edit
- The synchronous threshold makes "this is a bulk operation" a decision the platform makes, not one a user discovers

**What it costs**

- A window in which search results mix old and new parent values, which has to be explained to tenants rather than fixed
- Orchestration, checkpointing and a parent-to-children mapping to maintain — and a mapping that drifts produces an incomplete fan-out that only reconciliation will find
- Rate limiting means a large edit is slow by design, which will be experienced as the platform being slow
- Two concurrent fan-out jobs touching overlapping children need ordering, which is extra correctness work for a rare case

**Choose differently when.** If parent fields were not filterable or facetable, storing the parent once and joining at query time would be clearly better — no fan-out, no job, no inconsistency window. The decision follows directly from ADR-07, and if that one is revisited this one goes with it.

**Why it holds up over time.** "Bound and report the work a single change can create" is a property of any denormalised store. The orchestration technology is incidental; what must survive is the threshold, the checkpoint and the published completion target.

> **Lesson.** In a denormalised store, find the change that touches the most documents and design for it explicitly. It is never the average change, and it is always the one that takes the platform down.

#### ADR-09 · Idempotency lives in the data model: apply is guarded on (source, entity key, source version)

**Status:** Accepted  ·  **Shown on views:** 13, 12, 21

*Redelivery and reordering are certain. What stops a replayed change from overwriting a newer value?*

**Context.** Every durable transport on this path offers at-least-once delivery, and every partitioned transport preserves order only within a partition. A consumer restart replays, a producer retry duplicates, a rebalance reorders within a window, and a backfill re-emits deliberately. The usual response is to chase exactly-once semantics in the transport, which is expensive where it is available and unavailable where it matters. The cheaper and more durable response is to make duplicate and out-of-order delivery harmless in the data model, so that the transport's weaker promise is sufficient. Source systems already supply the necessary input: a monotonic version per entity, which is what the platform needs to tell a replay from a change.

**Decision.** Every change carries (source, entity key, source version). Apply compares the incoming source version against the version already applied for that entity: equal means a duplicate and is dropped, older means a reorder and is discarded, newer is applied. A push without a source version is explicitly the weakest supported contract — last-writer-wins by arrival — and is labelled as such on the integration.

**How it is realised on AWS.** The DynamoDB assembly-state table holds applied_version per (entity_key, source_name) and the apply is a conditional write on it, so the guard is enforced at the store rather than in process memory. The guard runs twice: at capture, to avoid logging a change already superseded, and at apply, because the log is partitioned and a consumer can see a reorder the capture tier did not. Fast-lane partial updates carry a per-field version so a field-level write cannot be reverted by a stale whole-document write from the slow lane.

| Option | Verdict | Reasoning |
|---|---|---|
| Version-guarded conditional apply in the data model | Chosen | Makes at-least-once sufficient and reordering harmless; costs one conditional write per change and a hard dependency on source versions |
| Transactional exactly-once in the transport | Rejected | Available in some engines, expensive, and still does not survive a backfill, a re-emit or a second producer path |
| Deduplication window on event ids | Rejected | Catches duplicates and not reorders, and its correctness is bounded by a window chosen before the outage that will exceed it |
| Last-writer-wins by arrival time everywhere | Rejected | Simple and silently wrong: a replay after a lag spike reverts a document to an older value with no error anywhere |

**What it buys**

- Redelivery becomes a comparison and reordering becomes a discard, which removes a whole class of correctness work from the pipeline
- Backfill, re-emit after reconciliation, and a full rebuild are all the same operation as a live change, so there is one code path and it is constantly exercised
- A reordered change produces a measurable counter rather than a silent regression, which is a signal worth alarming on
- Because the guard is at the store, a process restart or a split consumer cannot reopen the hole

**What it costs**

- A hard integration requirement: a source that cannot emit a monotonic version gets the weakest contract and a visibly worse guarantee
- One conditional write per change on the hot path, which at 25,000 changes a second is a real capacity line
- Field-level versions on the fast lane add bookkeeping to the cheapest path in the system
- The guard rejects legitimate corrections that arrive with a stale version, so a source that rewrites history needs an explicit re-emit with a new version

**Choose differently when.** If every source could emit a global transaction order and the platform consumed a single totally ordered stream, arrival order would be sufficient and the guard would be redundant. The decision is justified by four independent sources, partitioned transport and deliberate replay.

**Why it holds up over time.** "Put idempotency in the key, not the transport" has survived every generation of messaging technology, and it will survive this one. What would change it is a source model where versions do not exist — in which case the weaker contract is the honest answer, as it already is for the push path.

> **Lesson.** Do not buy exactly-once. Make duplicates and reorders harmless in the data model, and then let the transport be as weak as it honestly is.

### Evidence before promotion

*What makes a new index or a new ranking good enough to serve, and how a lost document is found.*

#### ADR-10 · The swap gate is four independent machine-checkable detectors, and it refuses rather than escalating

**Status:** Accepted  ·  **Shown on views:** 14, 18, 06

*A rebuild has finished. What has to be true before forty million people start seeing it, and who decides?*

**Context.** A completed rebuild is not a good index. It can be truncated because a source partition stalled, structurally wrong because a mapping change dropped a field, semantically worse because an analysis change hurt relevance, or broken in a way nobody predicted. Each of those is invisible to the others: a document count does not notice a changed analyser, a field diff does not notice a relevance regression, and an offline relevance score does not notice that half the catalogue is missing. The temptation is to collapse the check into a human sign-off, which is both the least reliable detector and the one most likely to be asked for at 2am, when a scheduled rebuild finishes and the person paged has no way to evaluate the question they have been handed.

**Decision.** Promotion is gated on four independent detectors: document count within a declared tolerance of the source count; zero unmapped-field errors; a sampled field-level diff against the live index; and a judgement-set relevance score no worse than live by more than a declared margin, with a canary traffic slice where one is configured. The gate refuses on failure. It does not escalate to a human, and there is no override path that skips it.

**How it is realised on AWS.** The reindex orchestrator runs the gate as a step in the Step Functions execution before the alias action, recording each detector's result in the reindex_job row. The diff detector relies on assembly determinism (ADR-07) to compare a sample of documents between candidate and live. Where no usable judgement set exists, the detector reports "no usable judgement set" rather than a pass, and the gate's verdict says which detectors actually ran. A refusal leaves the candidate index in place with its checkpoints, so the cause can be fixed and the build resumed rather than restarted.

| Option | Verdict | Reasoning |
|---|---|---|
| Four independent detectors, gate refuses | Chosen | Each catches what the others cannot, and the refusal is reliable at 2am; costs gate latency and a dependency on maintained judgement sets |
| Human sign-off on each swap | Rejected | The least reliable detector, needed most when judgement is worst, and it turns a routine operation into an interrupt |
| Document count only | Rejected | Cheap and blind to every structural and semantic failure; a mapping that drops a field passes it perfectly |
| Canary alone, promote and watch | Rejected | Catches the unpredicted and exposes real users to the predictable; better as the fourth detector than as the only one |

**What it buys**

- Four failure classes are caught mechanically, before any reader is exposed to the candidate index
- A refusal is a normal outcome with a cause attached, and the rebuild resumes from its checkpoints rather than starting again
- Because the gate reports which detectors ran, a weak check cannot masquerade as a strong one
- Nobody is asked to approve a swap they are not equipped to evaluate, which is what makes scheduled rebuilds genuinely unattended

**What it costs**

- A maintained judgement set per index becomes an operational obligation, and an unmaintained one makes the gate confidently wrong
- The diff detector binds the platform to deterministic assembly forever
- Gate latency is added to every rebuild, and a flaky detector blocks a correct index — a failure mode the gate itself creates
- A deliberate large change (a new analyser that legitimately changes most documents) will trip the diff detector, so the tolerance has to be declarable per promotion, which is a small override in all but name

**Choose differently when.** For a brand-new index with no judgement set, no live index to diff against and no traffic to canary, only the count and mapping detectors apply — and the gate should say so rather than pretend. If an index were small enough to verify exhaustively, sampling would be unnecessary and the gate could be exact.

**Why it holds up over time.** "Promotion requires evidence, and the gate may refuse" is the durable part. Better detectors will arrive — learned diffs, interleaving, counterfactual estimators — and should slot into the same gate. What must not drift is the escalation path, because a gate that escalates is a gate that will be overridden.

> **Lesson.** Build gates out of detectors that fail differently, and let the gate say no. A check that can only escalate is a notification.

#### ADR-11 · Relevance is versioned configuration inside the index definition, split explicitly into index-time and query-time classes

**Status:** Accepted  ·  **Shown on views:** 06, 14, 20

*An engineer wants dishes to rank above shop names. Is that a sixty-second configuration change or a four-hour rebuild — and when do they find out?*

**Context.** Relevance work is iterative and never finished, so the cost of each iteration sets how much of it happens. Some relevance changes are query-time: boosts, decay parameters, tie-breaks, business rules. Others are index-time: analysis, index-time synonyms, materialised scoring signals, anything baked into the tokens on disk. They look identical in a configuration file and differ by four hours of rebuild. If the platform does not distinguish them, engineers discover the difference after planning a release around the wrong answer, and — worse — the two get mixed in one change, so the fast rollback no longer covers what was actually shipped.

**Decision.** Relevance configuration lives inside the versioned index definition and is promoted and rolled back with it. The platform classifies every relevance change as query-time or index-time at authoring time and reports which it is. A query-time change takes effect, and rolls back, within 60 seconds with no rebuild. An index-time change is a breaking definition change and takes the full rebuild-and-swap path of ADR-03.

**How it is realised on AWS.** The Aurora definition registry holds relevance configuration as part of the definition version; the query service resolves the active relevance version per request from a short-TTL cache, which is what makes a 60-second rollback reach live traffic without a deploy. The definition API diffs a candidate against the current version and labels the change class before accepting it. Two variants can be active behind one alias with a declared traffic split, and the query and engagement log records the relevance version served so per-variant metrics are reconstructible.

| Option | Verdict | Reasoning |
|---|---|---|
| Relevance inside the versioned definition, classes split and reported | Chosen | One promotion and rollback mechanism for the whole definition, with honest cost signalling; costs a registry lookup on the query path |
| Relevance as separate runtime configuration outside the definition | Rejected | Faster to change and impossible to reason about: a stored index and the ranking applied to it drift into combinations nobody chose |
| Relevance in query-building code, deployed with the service | Rejected | Makes every ranking experiment a service release and couples relevance iteration speed to deployment cadence |
| One class only — treat every relevance change as index-time | Rejected | Safe and far too slow: it taxes the cheap changes at the price of the expensive ones and suppresses iteration |

**What it buys**

- An engineer learns the cost of a change before scheduling it, which is the difference between a weekly cadence and a quarterly one
- A ranking regression is reversible in a minute for the common case, without a rebuild and without a deploy
- Because the relevance version is recorded per query, a win or a loss can be attributed to a specific configuration afterwards
- Index and ranking are versioned together, so there is no combination of the two that the registry cannot name

**What it costs**

- The registry becomes a query-path dependency; its cache TTL silently bounds how fast a rollback actually propagates
- A mixed change containing both classes must be split or treated as index-time, which is a rule engineers will meet at an inconvenient moment
- Two variants behind one alias need enough traffic per variant to resolve a difference, which small tenants will not have
- Versioning relevance with the definition means a trivial boost change creates a new definition version, and the version history gets noisy

**Choose differently when.** If the engine could re-analyse documents in place cheaply, the index-time class would shrink to almost nothing and the distinction would stop mattering. If relevance were learned and served by a model outside the index, this decision would move to that model's deployment story instead.

**Why it holds up over time.** "Configuration that changes stored bytes is a different class from configuration that changes a query" is a property of search engines generally, not of one engine. The classification lists will change; the split will not.

> **Lesson.** When two changes look the same and cost differently by orders of magnitude, make the platform say which one the user just asked for — before they commit to it.

#### ADR-12 · A silently lost document is found by continuous sampled reconciliation, and a silent source by a quiet-period alarm

**Status:** Accepted  ·  **Shown on views:** 17, 21, 18

*A change never reached the platform. What notices, and how long does it take?*

**Context.** This is the failure the platform is least equipped to see, because it has no symptom. A dropped change produces a document that is absent or stale, with no error, no retry, no dead letter and no alert — and the first report arrives weeks later from a merchant asking why their listing is invisible, by which time nobody can reconstruct what happened. Every other failure class in this package announces itself. This one has to be hunted. The available instruments are: compare samples against the source continuously; compare counts or checksums per index periodically; rebuild often enough that any loss is bounded by the rebuild interval; or require source-side sequence numbers and detect gaps in the log. They cost very different amounts and bound the error very differently.

**Decision.** A reconciler continuously samples documents, compares them field by field against their sources, and publishes a divergence rate as a first-class signal with a target of ≤ 0.01%. Separately, every source binding declares a quiet period, and emitting nothing for longer than it raises an alarm — an absence of signal is treated as a signal. A detected divergence triggers a targeted re-emit for the affected keys from the snapshot rather than a full rebuild.

**How it is realised on AWS.** An ECS reconciler samples 1,000,000 documents a day, weighted towards recently changed entities and fast-lane fields, reading the source through its own read path and the index through the alias. Divergences are recorded with the field, the index version and both values, and re-emitted into the change log with a fresh source version so the normal apply path repairs them. Quiet-period alarms are configured per source binding in the registry, and the schema-drift detector shares the same comparison output.

| Option | Verdict | Reasoning |
|---|---|---|
| Continuous sampled reconciliation plus per-binding quiet-period alarms | Chosen | Bounds the error with a published number and finds drift as a side effect; costs source read load and a divergence bound that is a function of sample rate |
| Periodic full count-and-checksum per index | Rejected | Exact and expensive, and only as good as the sources' ability to produce a comparable digest — which two of the four cannot |
| Frequent scheduled full rebuilds as the detector | Rejected | Bounds the loss by the rebuild interval and pays four hours of cluster capacity to learn something sampling learns continuously |
| Source-side sequence numbers with gap detection in the log | Rejected | The cheapest and strongest detector where it is available; rejected as the primary mechanism only because it cannot be required of every source, and retained as a per-source enhancement |

**What it buys**

- The platform's worst failure acquires a number, a trend and an alarm, instead of a support ticket
- Repair is a targeted re-emit through the normal path, so there is no bespoke correction code to be wrong
- The same comparison detects schema drift and fast-lane field reverts, which are two other silent classes
- A quiet-period alarm per binding distinguishes "nothing happened" from "nothing arrived", which no rate threshold can

**What it costs**

- The divergence bound is a function of sample rate: at a million documents a day against eighty million, a single lost document can stay undetected for weeks, and only the aggregate rate is really bounded
- Sampling puts read load on the sources, which is exactly what source owners were promised the platform would avoid
- Weighting the sample towards recent changes makes the long tail of old, never-touched documents the least-checked part of the index
- Quiet periods are per-binding guesses, and one set too generously hides an outage for as long as it lasts

**Choose differently when.** If every source could emit a cheap comparable digest, a periodic full comparison would replace sampling and bound the error exactly. That is the preferred end state, and it is a per-source negotiation rather than a platform decision.

**Why it holds up over time.** An absent change will still produce no error in a decade. Sampling against the source, and alarming on absence rather than on a rate, remains the only general detector whatever the pipeline is built from.

> **Lesson.** For every failure that produces no error, name the instrument that will find it and publish what that instrument can actually bound. "We would notice" is not an instrument.

### Boundaries and placement

*What a query is allowed to see, who may move an alias, and where the platform runs.*

#### ADR-13 · Every volatile and access-controlling field is filtered before retrieval, from the caller's identity, never from the request

**Status:** Accepted  ·  **Shown on views:** 20, 19, 04

*Where are unavailable, suppressed and unauthorised documents removed — before the engine scores, or after it answers?*

**Context.** Post-filtering is the path of least resistance: let the engine return results, drop the ones the caller should not have, send the rest. It is also visibly broken in three ways users notice immediately. A page of ten becomes a page of six. Facet counts disagree with the results beneath them. Deep pagination skips and repeats, because the offset was computed over a candidate set that included documents the caller never saw. The access case is worse than cosmetic: a filter the caller supplies is a filter the caller can omit, and a filter applied after retrieval has already let the engine score and count documents the caller is not entitled to.

**Decision.** Every field in the fast freshness class and every access-controlling field is indexed and applied as a mandatory pre-retrieval filter. Access filters are derived at the query service from the validated identity, never accepted as a request parameter. The platform does not rely on callers post-filtering, and a query that would require a post-filter for correctness is refused.

**How it is realised on AWS.** The query service resolves the caller's principal and tenant from the validated token, derives the permitted-visibility clause, and composes it into the engine query as a required filter before any scoring clause. Fast-lane fields — availability, stock, open/closed, moderation state — are present in the document (ADR-04, ADR-07), so they filter natively. Cursor pagination replaces offset beyond a declared depth, and a cursor is valid only against the index version that issued it. Over-ceiling queries get a typed rejection naming the limit rather than a degraded answer.

| Option | Verdict | Reasoning |
|---|---|---|
| Mandatory pre-retrieval filtering, access derived from identity | Chosen | Correct counts, correct pagination, no omittable filter; requires every filterable field to be in the index, which constrains ADR-04 and ADR-07 |
| Post-filter in the query service | Rejected | Breaks page sizes and facet counts, and lets the engine score documents the caller is not entitled to |
| Post-filter in the calling application | Rejected | The same failure, distributed across every client, with each client's bug becoming the platform's incident |
| Caller-supplied visibility filter, validated by the platform | Rejected | Validation of an optional parameter is not the same as a mandatory clause, and the failure mode is a disclosure |

**What it buys**

- Facet counts, page sizes and pagination are consistent with what the caller can actually see
- Access control cannot be bypassed by omitting a parameter, and no unauthorised document is scored
- Because volatile fields are present in the index, "open now" and "in stock" are real filters rather than post-hoc trimming
- A cost ceiling can be enforced before execution, since the full query shape including filters is known up front

**What it costs**

- Every field that must be filterable has to be indexed and kept fresh, which is write cost and the reason the fast lane exists
- Access constraints in the document mean a permission change is an indexing event, with the staleness that implies — so suppression is fail-closed and fast lane (ADR-05)
- Cursor-based deep pagination is a harder client contract than offsets, and cursors break across a swap
- Highly dynamic per-principal permissions would make the document's access fields churn, which is a limit on how far this scales

**Choose differently when.** If permissions were so dynamic and principal-specific that indexing them were impossible — a document-level ACL per user in a large organisation — the design would have to move to per-principal partitioning or accept a post-filter with its pagination cost, declared honestly. The decision assumes coarse, slow-moving visibility classes.

**Why it holds up over time.** "Filter in the engine, derive from identity" is a security and correctness property that no engine change affects. What would change is how permissions are represented, which is why they are indexed fields rather than hard-coded clauses.

> **Lesson.** A filter applied after retrieval is not a filter; it is a redaction with broken arithmetic. And a filter the caller supplies is a filter the caller can forget.

#### ADR-14 · Moving an alias is a privilege granted separately from authoring a definition

**Status:** Accepted  ·  **Shown on views:** 19, 14, 09

*Which single action in this platform changes what forty million people see, instantly, with no deploy and no review?*

**Context.** The answer is the alias move, and it does not look dangerous. It is one API call, it is fast, it is reversible, and it is performed by operations people doing routine work. That combination is exactly why it is the most under-protected privilege in a platform of this shape: authoring a definition looks like the consequential act because it involves thinking, while the swap looks like plumbing. But a definition is inert until an index is built from it and the alias points at it. Everything upstream — validation, classification, dry run, gate — exists to make the swap safe, and all of it is bypassed by someone who can move an alias directly.

**Decision.** Four privileges are separate grants: query an index, write to a source, change an index definition, and move or roll back an alias. The alias privilege is the most restricted, is never exposed as an integration surface, and every use is written to an append-only audit store with the actor, the previous and new index versions, and a reason. The orchestrator's own alias action is itself a distinct identity with no human members.

**How it is realised on AWS.** Alias actions are performed only by the alias manager, assumed by two principals: the reindex orchestrator's role after a passing gate, and a break-glass human role with a short-lived session, approval and a mandatory reason field. Both write to the S3 Object Lock audit store alongside the registry's promotion history. The definition API cannot move an alias; the query path has no write permission of any kind; and the admin console's routine role can request a rebuild but not promote its result.

| Option | Verdict | Reasoning |
|---|---|---|
| Alias move as its own restricted grant, audited with a reason | Chosen | Matches the privilege to the blast radius; costs an extra grant to administer and a break-glass path to keep working |
| One admin role covering definitions and aliases | Rejected | Makes the most consequential action in the platform a side effect of ordinary administrative access |
| Alias move only ever performed by the orchestrator | Rejected | Removes the human blast radius and also removes the rollback path when the orchestrator itself is the problem; the break-glass role exists for exactly that |
| Alias move requires two-person approval always | Rejected | Correct for a promotion and wrong for a rollback: it puts a second human between the platform and a one-minute recovery |

**What it buys**

- The action with the largest blast radius has the smallest set of principals and the strongest record
- The gate cannot be bypassed accidentally, because the role that could bypass it is not the role anyone holds day to day
- A reason field attached to every swap makes the promotion history readable a year later, which is what makes incident review possible
- Rollback stays a one-minute operation, because the break-glass path is designed for it rather than discovered during it

**What it costs**

- A break-glass path is a standing risk that has to be exercised to stay working, and exercising it is itself a privileged action
- More grants to administer, and an on-call engineer who cannot promote an index they have correctly judged ready
- A reason field is only as good as what people type into it under pressure
- Separating request-a-rebuild from promote-its-result means routine work needs two roles, which invites a shortcut nobody will document

**Choose differently when.** In a single-team platform with one index and five engineers, this separation is ceremony: everyone holds every grant anyway and the audit trail is the chat log. It earns its cost at twelve tenants, several teams, and an index that forty million people search.

**Why it holds up over time.** "Grant the privilege that matches the blast radius, not the one that matches the job title" is an access-control principle that outlives any identity system. What changes is the mechanism, not the boundary.

> **Lesson.** Find the single call in your architecture that changes what everyone sees, and protect that one specifically. It is rarely the one that looks important.

#### ADR-15 · One write region; read regions serve a follower index that is stale and says so

**Status:** Accepted  ·  **Shown on views:** 16, 11, 21

*Can two regions index the same entity, and if not, what does a region loss cost?*

**Context.** Multi-region is usually argued as an availability improvement, and for a read path it is. For this write path it is a correctness question, because assembly is stateful: the applied source version, the join partners, the parent-to-children mapping and the enrichment cache are all per-entity state. Two regions assembling the same entity independently produce two documents from two views of that state, and there is no alias, version guard or merge rule in this design that reconciles them — the version guard assumes one writer per entity. The alternative is a single write region, which concentrates the write path's availability risk and makes a regional failover an indexing pause rather than a seamless continuation.

**Decision.** One write region owns capture, assembly, index writing and the registry. Read regions serve a follower index with a declared staleness and take no local writes. On losing the write region, reads continue in the read regions, stale and declared; indexing resumes from the change log when the write region or its replacement is available, with an RTO of 15 minutes for the query path.

**How it is realised on AWS.** eu-west-1 is the write region across three availability zones: MSK brokers, assembly and writer tasks, the OpenSearch domain's data nodes, and the Aurora registry with multi-AZ and point-in-time recovery. Cross-region replication ships index segments to follower domains in the read regions, whose lag is published and stamped on responses served from them. Route 53 health checks shift query traffic to a read region on failure; no read region holds an alias-move privilege or an assembly-state table, so there is nothing there to diverge.

| Option | Verdict | Reasoning |
|---|---|---|
| Single write region, follower read regions, declared staleness | Chosen | Keeps per-entity write ownership unambiguous and read availability high; costs an indexing pause during a regional failover |
| Active-active indexing in two regions | Rejected | Two writers per entity with no merge rule; it would require entity ownership partitioning and a reconciliation protocol this topic does not justify |
| Single region, no read regions | Rejected | Simplest and leaves a 15,000-queries-a-second read path with no answer to a regional failure |
| Entity-partitioned write ownership across regions | Rejected | Correct and genuinely available, and it imports a partition-ownership and failover protocol that is a larger system than this one |

**What it buys**

- Per-entity write ownership is unambiguous, which is what makes the version guard in ADR-09 sound
- The read path survives a regional failure with a declared staleness instead of an outage
- Because read regions hold no write state, failover has nothing to reconcile and failback has nothing to merge
- Indexing resumes from the log rather than from a follower, so a failover cannot silently lose changes

**What it costs**

- A write-region outage is an indexing pause: freshness degrades for everyone until it is restored, and the response has to say so
- Follower lag is a second staleness dimension on top of lane freshness, and callers now need to understand both
- Read-region capacity is paid for continuously and used fully only during an incident
- The write region is a concentration of risk that no amount of multi-AZ removes

**Choose differently when.** If the platform's documents were strictly partitionable by region — a marketplace whose listings never cross borders — regional write ownership would be natural and active-active would follow cheaply. If the freshness obligation during a regional outage became intolerable, the entity-partitioned option returns, and with it a far larger design.

**Why it holds up over time.** "Stateful per-entity assembly needs one writer" is a consequence of the version-guard model, not of a cloud's topology. It holds until the data model gains a merge rule, which would be a different architecture rather than a different deployment.

> **Lesson.** Decide multi-region on the write path's state model, not on the read path's availability. If two writers cannot be reconciled by a rule you can state in a sentence, you have one write region.

#### ADR-16 · Shared cluster with per-tenant indices by default; dedicated clusters by declared policy, with live migration

**Status:** Accepted  ·  **Shown on views:** 08, 16, 21

*Does every tenant share one search cluster, and what happens the first time one of them cannot?*

**Context.** Twelve tenants of very different sizes share a platform whose dominant costs are cluster capacity and unused headroom. One cluster with per-tenant indices gives the best utilisation and the worst blast radius: a hostile query, a catalogue reload or a rebuild in one tenant competes for the same nodes as another tenant's p99. A cluster per tenant inverts that — small clusters idling expensively, twelve upgrade paths, twelve sets of headroom — and buys isolation that most tenants do not need. The decision is harder than a utilisation calculation, because for some tenants isolation is not about performance at all: residency obligations or key separation can make sharing a cluster a compliance question, and no quota answers a compliance question.

**Decision.** The default is a shared cluster with per-tenant indices, per-tenant quotas on documents, storage, indexing rate, reindex concurrency and query rate, and resource isolation between query and indexing traffic. A declared placement policy — driven by tenant size, noisy-neighbour risk, and residency or key-isolation obligations — moves a tenant to a dedicated cluster, and migration between the two is a supported operation requiring no change to the tenant's definition or its callers.

**How it is realised on AWS.** Per-tenant indices on the shared OpenSearch domain, with quotas enforced at the query service and the capture tier and rejections typed and attributed. Query and indexing traffic are isolated at the resource level, and the reindex writer is confined to a reserved share. Placement is a field on the tenant record in the registry; migration is a rebuild into a dedicated domain followed by an alias move, which is to say it is the mechanism of ADR-01 and ADR-03 reused rather than a new path.

| Option | Verdict | Reasoning |
|---|---|---|
| Shared by default, dedicated by declared policy, live migration | Chosen | Best utilisation for the many, a real escape for the few, and migration built from mechanisms that already exist |
| Shared cluster for everyone, quotas only | Rejected | Has no answer to a residency obligation, and a quota cannot fully protect a p99 from a co-tenant's rebuild |
| Dedicated cluster per tenant | Rejected | Twelve sets of idle headroom and twelve upgrade paths to buy isolation most tenants do not need |
| Shared clusters grouped by tenant tier | Rejected | A reasonable middle that adds a placement dimension without removing the compliance case; revisit if the tenant count grows past a few dozen |

**What it buys**

- Utilisation stays high for the majority while the minority with a hard requirement has a defined route out
- Migration reuses rebuild-and-swap, so the escape hatch is exercised by the same code path used every week
- Quotas and typed rejections make noisy-neighbour effects attributable rather than mysterious
- Placement is data in the registry, so the policy is reviewable rather than tribal

**What it costs**

- A shared cluster means a shared upgrade, a shared version and a shared failure domain, whatever the quotas say
- Reserved shares for reindex and query mean paying for headroom that is idle most of the time
- Migration is hours of rebuild per tenant, so "live" means uninterrupted rather than instant
- A policy written after the first dedicated tenant arrives will be a rationalisation of a decision already made under pressure

**Choose differently when.** At a few dozen tenants, tiered shared clusters become the better structure and this decision is superseded by a placement dimension rather than a binary. At a single large tenant, the whole question disappears and the cluster is simply sized for it.

**Why it holds up over time.** "Share by default, isolate by declared policy, and make the move a supported path" is a multi-tenancy posture rather than a technology choice. What changes is the trigger list, which is why it is policy data rather than code.

> **Lesson.** Write the placement policy before the second large tenant onboards. Isolation demanded for compliance cannot be answered with a quota, and a migration path invented during an escalation is not a path.

## Every package used, in one table

These terms are used precisely in this package. Several are used loosely in the wider literature on search platforms, and the difference matters when reading the decision records.

| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Alias | The only reader-facing name for a logical index. It points at exactly one index version at a time, or at two during a declared relevance variant split. | The architecture's central contract: the thing a reader holds so that the index behind it can be replaced. | "Index name", which in this package always means the real, versioned, private name a reader never sees. |
| Index version | One physical index built from one definition version, named with that version (idx_listings_v41), in state building, live or retained. | The unit of rebuild, gating, swap and rollback. | "The index", which in casual use conflates the alias, the version and the definition. |
| Index definition | The immutable, versioned declaration of a document's shape, sources and join keys, per-field analysis, freshness classes, retention, replicas and relevance configuration. | The only way to create or change an index; the input a rebuild is reproducible from. | "Mapping", which is only the engine-facing part of it. |
| Freshness class | A declared budget attached to a field group — fast lane, suppression or slow lane — with its own queue, writer and end-to-end target. | The mechanism that spends freshness where a user can catch us being wrong. | "Refresh interval", which is an engine setting and only one contributor to this number. |
| Observed freshness | The measured source-commit-to-searchable lag of the index actually served, stamped on the response alongside the index version. | What makes a stale answer honest rather than misleading, and a staleness complaint reproducible. | "Indexing lag", usually meaning consumer lag, which can read zero while the index is wrong. |
| Acceptance gate | Four independent machine-checkable detectors — document count, mapping errors, sampled field diff, judgement-set score, with an optional canary — run before an alias move. | The reason a bad index is refused rather than discovered. | "Smoke test", which implies one check and a human reading the result. |
| Fan-out job | The tracked, rate-limited, resumable rebuilding of the child documents affected by a change to a shared parent entity. | Bounds what a single edit can cost and gives it a completion target. | "Bulk update", which does not imply a bound, a checkpoint or a progress surface. |
| Divergence rate | The share of sampled documents found to differ from their sources by continuous reconciliation, published as a first-class signal. | The only number that says whether the index still matches the world, and the only detector of a silently lost document. | "Index health", which usually means shard and node status and says nothing about correctness. |
| Owner overlay | A read-path merge of the caller's own pending changes, read from a short-TTL recent-writes table, applied only to entities that caller owns. | Delivers one-second read-your-writes without imposing it on the index's freshness budget. | "Read-your-writes consistency", which in most systems is a property of the store rather than of one principal's read path. |
| Pre-retrieval filter | A filter composed into the engine query before scoring, derived from the caller's identity for access fields and from indexed values for volatile ones. | Keeps facet counts, page sizes and pagination consistent with what the caller can actually have. | "Post-filter", the thing this package forbids for suppression and access. |
