# Search Indexing Service

**Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-10 · 21 views · 16 architecture decision records**

You search a delivery app at 8pm and the first restaurant is closed. You search a marketplace for earbuds, tap the third result, and it is out of stock. You rename a file in a workspace tool and the old name keeps coming back for a minute. None of those are ranking failures — the cluster answered correctly from a copy of the world that was wrong. This package is the platform that keeps that copy right: a multi-tenant search indexing service inside one consumer marketplace — an assumed 40 million monthly active users, 180,000 active merchants, 12 product tenants, 80 million searchable documents across 40 indices, 25,000 change events a second rising to 100,000 in a burst, 15,000 queries a second rising to 45,000 at the dinner hour, and a largest index of 38 million documents that has to rebuild inside four hours — built on a managed durable change log, a managed search cluster, strongly consistent assembly state, and a relational registry holding every index definition the platform has ever served.

The product is not the ability to search. It is **the ability to be replaced**: a search index accumulates reasons to be rebuilt faster than almost any other store, and the only question that matters is whether replacing it is a routine operation or an incident.

The design rests on one rule: **readers address an alias, never an index — every index is a disposable, versioned projection built from a retained change log, and the alias is the only contract.**

The decisions that carry the design:

- **Readers hold the alias, never the index.** One indirection, installed before anyone needs it, turns a breaking schema change, a wrong analyser, a corrupt shard and a bad relevance deploy from four separate incidents into one routine operation with one rollback (ADR-01).
- **The change log is the platform's own system of record for what changed.** Sources stay authoritative for what is true. Retention stops being a backup policy and becomes a replay-throughput commitment measured against the four-hour RTO (ADR-02).
- **Rebuild beside, then swap.** Double storage for the build window and a full replay buy an atomic cutover, a free rollback and somewhere to stand while validating. In-place mapping evolution is a one-way door (ADR-03).
- **Freshness is differentiated, not averaged.** Two lanes with separate queues, because a price on a popular item changes hundreds of times an hour and its description does not — and the fields a user catches instantly are also the cheap ones to write (ADR-04).
- **Absence beats a wrong presence.** Suppression, deletion and access filters fail closed with no tolerated staleness; enrichment values and ranking signals fail open, degraded and declared. A swap cannot save us from showing something that should already be gone (ADR-05).
- **One merchant's expectation is not a platform requirement.** Read-your-writes for the entity's owner is a read-path overlay on a short-TTL table, not a sub-second freshness obligation on eighty million documents (ADR-06).
- **Documents are assembled whole and say what they were built from.** Flat documents keep the query inside a 250 ms p99 and make volatile fields filterable; recorded source versions are what make an index reconcilable rather than a rumour (ADR-07).
- **A brand rename is a job, not a side effect.** Two million child rewrites are rate-limited, checkpointed, preemptible and reportable, with a bound on what one change may enqueue synchronously — and a mid-fan-out inconsistency that is published rather than hidden (ADR-08).
- **Idempotency lives in the data model.** Guarded on (source, entity key, source version) at the store, redelivery becomes a comparison and reordering a discard, which is what lets at-least-once transport be sufficient (ADR-09).
- **The gate refuses; it does not escalate.** Four independent detectors, because document count catches a truncated rebuild, sampled diff catches a mapping change, judgement score catches a relevance regression, and a canary catches what none of them predicted. Nobody is asked at 2am to approve a swap they cannot evaluate (ADR-10).
- **The platform says what a relevance change costs before it is scheduled.** Query-time changes roll back in sixty seconds; index-time changes are breaking and take the rebuild path — and an engineer finds out which before planning a release around the wrong answer (ADR-11).
- **The lost document is hunted, not reported.** Continuous sampled reconciliation publishes a divergence rate, and a quiet-period alarm per source binding turns an absence of signal into a signal (ADR-12).
- **Filter before retrieval, from identity.** A post-filter breaks pagination and facet counts in exactly the way users notice, and a caller-supplied filter is one a caller can omit (ADR-13).
- **Moving an alias is its own privilege.** It is the one call that changes what forty million people see, instantly, with no deploy — so it is granted separately from authoring a definition and audited with the actor and the reason (ADR-14).
- **One write region, because two writers cannot be reconciled by any rule in this design.** Read regions serve a follower index that is stale and says so; indexing resumes from the log rather than from a follower (ADR-15).
- **Share by default, isolate by declared policy.** Placement is registry data driven by tenant size, noisy-neighbour risk and residency — and migration reuses rebuild-and-swap rather than inventing a path during an escalation (ADR-16).

The architecture one-pager (including why the design should still hold up in ten years, and the six risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~16 min) and [docs/decision-record.md](docs/decision-record.md) (~68 min).

---

## What is here

| Path | Contents |
|---|---|
| `diagrams/index.html` | The landing page: 21 views in seven acts with every format linked, then the **architecture one-pager** and the **decision record** |
| `diagrams/*.html` | One self-contained page per view: the inlined diagram plus the reasoning cards, with copy / PNG / PDF export |
| `diagrams/svg/*.svg` | The same views as SVG with the diagram XML embedded; they re-open fully editable in diagrams.net |
| `diagrams/drawio/*.drawio` | draw.io source |
| `docs/architecture-one-pager.md` | The one-pager as markdown |
| `docs/decision-record.md` | The 16 decision records, the capability-to-technology table and the package glossary as markdown |
| `specs/part-a..c.json` | Diagram specifications, the source of truth for every view |
| `specs/manifest-a..b.json` | Acts, page titles, subtitles and reasoning cards |
| `specs/adr-onepager.json`, `specs/adr-records-a..b.json` | The one-pager, the decision records, the capability-to-technology table and the glossary |
| `scripts/build.sh` | Rebuilds every deliverable from the specs (Node only, no network) |
| `ask.md` | The requirement |

## The twenty-one views

| # | View | What it answers |
|---|---|---|
| 01 | System Context | Who searches, who publishes, and the fact that the platform holds no authoritative copy of anything |
| 02 | High-Level Architecture | Seven stages, with the alias and its gate deliberately sitting between the index and the reader |
| 03 | Actors and Their Core Journeys | Six parties, two of them machines — including the reconciler, which is in the cast because the worst failure has no complainant |
| 04 | Journey — A Consumer Finds Something That Exists | The trough is the tap, one screen after a search that looked fine |
| 05 | Journey — A Merchant Publishes a Change | The one user who checks immediately, and why they do not get to set the budget for 80 million documents |
| 06 | Journey — A Search Engineer Ships a Relevance Change | The trough is promotion: finding out that what you wanted is a four-hour rebuild |
| 07 | Layered Architecture | Eight layers and one seam: above the log is what changed, below it is a projection |
| 08 | Platform Components | Four planes in one region, and the two components allowed to face outwards |
| 09 | Integration Surface | Four ways in, four ways out, three authorisation stories — and the swap privilege is not among them |
| 10 | Data Flow — One Change, Seven States | Only two of the seven are authoritative |
| 11 | Storage Zones | Why the search indices are deliberately the one zone that is not backed up |
| 12 | Data Model | Twelve entities, and one composite key doing the work of a deduplication service |
| 13 | Critical Flow — A Price Change Becomes Searchable | Fifteen messages, and the acknowledgement that arrives before the document is searchable |
| 14 | Reindex and Swap | Build beside, catch up, gate on four detectors, move one alias |
| 15 | Freshness Lanes | Five write classes, five budgets, and no shared queue between them |
| 16 | Deployment Architecture | One write region, and read regions that are honest about being behind |
| 17 | Observability | Signal by stage, plus one row that exists because the others cannot see this platform's worst failure |
| 18 | Index Lifecycle | Seven stages, and the two places the loop is allowed to say no |
| 19 | Security Trust Zones | Six zones, and four privileges that are deliberately not the same privilege |
| 20 | Identity and Access — One Filtered Query | Fourteen messages, and the filter that is derived rather than supplied |
| 21 | Failure Classes and Their Answers | Ten classes with their structural answer and their accepted residual, plus the two that would change the design |

## Rebuilding

```bash
bash scripts/build.sh
```

Node 20+ and nothing else. No draw.io Desktop, no browser, no network. The build assembles
`specs/views.json` and `specs/manifest.json` from the authoring parts, generates the draw.io
sources, fails on any geometry or routing defect, renders the editable and plain SVGs, writes the
HTML pages and the index, injects the one-pager and the decision record, and proves every relative
link in `diagrams/` resolves.

At v1.0 the geometry gate reports 0 errors and 0 warnings across 21 files, the routing gate
0 errors and 6 clutter warnings across 165 edges, all 324 nodes resolve to an icon with no weak
matches, and the link check passes on 166 relative links. The 6 warnings are label stacking on the
system-context and integration views, where several relationship edges converge on one centre; they
were reduced from 12 by shortening labels and merging two inbound surfaces, and the remainder is
inherent to those two layouts.

## A note on the numbers

Every rate, latency, ratio, threshold and retention figure in this package is a **stated
assumption**, chosen to be defensible and arguable rather than measured. They are stated precisely
so that a reviewer can disagree with one and follow it to the decision that depends on it.
`ask.md` marks them as assumptions section by section; the decision record's evidence note
restates the operating context in one paragraph.
