Architecture One-Pager
Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-10
Search Indexing Service · Solution Architecture v1.0 · Amazon Web Services · Data Platform Architecture · 2026-10
Readers address an alias, never an index: every search index is a disposable, versioned projection built from a retained change log, and the alias is the only contract.
Everyone has met this system from the outside and misdiagnosed it. You search a delivery app at 8pm and the first restaurant is closed. You search a marketplace for earbuds, tap the third result, and it is out of stock. You rename a file in a workspace tool and the old name keeps coming back for a minute. None of those are ranking failures. The search cluster answered correctly from a copy of the world that was wrong — and the copy was wrong because keeping it right is a different and harder problem than querying it. That problem is a multi-tenant indexing platform: it has to turn change streams from systems of record it does not own into whole searchable documents, keep the fields a user can instantly catch us about fresh within seconds while the rest can wait a minute, absorb a brand rename that touches two million documents, rebuild a 38-million-document index inside four hours without the query path noticing, let a relevance change be measured before it ships and reversed in a minute, and find the one document that silently never arrived — a failure with no error, no alert and no complainant until a merchant asks why their listing is invisible.
Separate what changed from what is derived from it, and make every derived thing disposable. Changes are captured from source change streams (and, for sources that cannot emit one, a signed push API), deduplicated and version-guarded, and committed to a durable partitioned change log before the source is acknowledged — so the accept path owes the caller durability and nothing else. The log, plus a daily source snapshot in object storage, is the rebuild base. Assembly workers join declared sources into whole self-describing documents, resolve a shared-parent change into a rate-limited resumable fan-out job, and call enrichment with a timeout and a previous-value fallback so no document is ever withheld for an enrichment failure. Writers split by freshness class: a fast lane performing partial updates for stock, price and opening hours; a suppression lane with no tolerated staleness; a slow lane writing whole documents; and a throttled reindex writer confined to a reserved share of cluster capacity. Indices are versioned and disposable — idx_v41 serving, idx_v42 building, idx_v40 retained for 72 hours — and readers address only an alias. A new index becomes live by moving that alias, and only after a gate of four independent machine-checkable detectors agrees; the gate refuses rather than escalating. On the read side the query service resolves alias and relevance version per request, derives a mandatory pre-retrieval filter from the caller's identity, enforces a per-tenant cost ceiling, overlays the caller's own pending writes, and stamps every response with the index version and observed freshness it was served at. A reconciler samples a million documents a day against their sources and publishes a divergence rate, because that is the only mechanism that sees a document the pipeline lost.
What it is, and what it is not
- A disposable, versioned projection of a retained change log — not A durable store that happens to be searchable
- An alias contract that hides index identity from every reader — not A cluster whose index names leak into client code
- Differentiated freshness, budgeted per field group and declared on the response — not One freshness SLA averaged across everything in the document
- Deletion and suppression with no tolerated staleness — not A filter the caller is trusted to apply after retrieval
- A rebuild that is routine, rehearsed, checkpointed and resumable — not A reindex that is an incident response with a progress bar
- A promotion gated by evidence and reversed by an alias move — not A relevance change validated by the person who wrote it
- A divergence rate published continuously — not A trust that the pipeline delivered everything it was sent
- The platform's own system of record for what changed — not A system of record for any business entity
- Lexical and structured search over documents the platform assembles — not Semantic retrieval, embeddings, or a ranking service
The decisions that are the architecture
- Readers address an alias, never an index (ADR-01) — The alias is the only reader-facing name, which is the single indirection that makes rebuild, rollback, schema change and relevance experiments the same routine operation instead of four different incidents.
- The change log is the system of record for what changed (ADR-02) — Retention stops being a backup policy and becomes a replay-throughput commitment measured against the four-hour RTO: the log must reach back far enough, and replay fast enough, to rebuild the largest index.
- Rebuild beside, then swap — never update in place (ADR-03) — Double storage for the build window and a full replay buy an atomic cutover, a free rollback, a safe schema change and somewhere to stand while validating. In-place mapping evolution is a one-way door.
- Freshness is differentiated, not averaged (ADR-04) — Two lanes with separate queues and separate budgets, with partial document update on the fast lane, because a price on a popular item changes hundreds of times an hour and its description does not.
- Absence beats a wrong presence (ADR-05) — Suppression, deletion and access filters fail closed with no tolerated staleness; enrichment values and ranking signals fail open, degraded and declared. A swap cannot save us from showing something that should be gone.
- Read-your-writes is a read-path overlay (ADR-06) — The one merchant who checks their own listing sets a one-second requirement for themselves, not the freshness budget for eighty million documents.
- Assemble whole, self-describing documents on write (ADR-07) — A flat document makes the query trivial and the shared-parent change expensive; the document carries the source versions it was built from, because an unreconcilable index is a rumour.
- Parent fan-out is a tracked job, not a side effect (ADR-08) — A brand rename touching two million listings is rate-limited, checkpointed and reportable, with a bound on how much work one change may enqueue synchronously.
- Idempotency lives in the data model (ADR-09) — Keyed on (source, entity, source version), redelivery becomes a comparison and reordering becomes a discard — so at-least-once transport is a property to exploit rather than a problem to solve.
- The swap gate refuses; it does not escalate (ADR-10) — Four independent detectors, because count catches a truncated rebuild, sampled diff catches a mapping change, judgement score catches a relevance regression, and a canary catches what none of them predicted.
- Relevance is versioned configuration inside the definition (ADR-11) — Split explicitly into index-time and query-time classes, so the platform can tell an engineer whether what they asked for is a sixty-second rollback or a four-hour rebuild — before they plan a release around the wrong answer.
- The lost document is found by sampling, not by errors (ADR-12) — Continuous reconciliation publishes a divergence rate, and a quiet-period alarm per source binding turns an absence of signal into a signal.
- Filter before retrieval, from identity (ADR-13) — A post-filter breaks pagination and facet counts in exactly the way users notice, and a caller-supplied filter is one a caller can omit.
- Moving an alias is its own privilege (ADR-14) — It changes what forty million people see in one call with no deploy, so it is granted separately from authoring a definition and audited with the actor and the reason.
- One write region; read regions stale and honest (ADR-15) — Two regions assembling the same entity produce two documents that no alias reconciles. Read regions serve a follower index with a declared staleness and take no local writes.
- Shared cluster by default, dedicated by declared policy (ADR-16) — Placement is a policy driven by tenant size and noisy-neighbour risk, with live migration between the two, because for some tenants sharing a cluster is a compliance question rather than a performance one.
Why this should still be right in ten years
The search engine, the log, the container runtime and the cloud will all be replaced inside this platform's life. These are the properties that should outlast them.
- The alias is an indirection, not a product feature. Every search engine worth using has some form of pointer that readers can hold while the thing behind it is replaced — an alias, a view, a symlink, a routing rule. The decision is that readers hold the pointer and never the index. That survives a migration from one engine to another, and it is the one thing that makes such a migration possible at all.
- Replayability is a capability, not a storage choice. "The log is what a rebuild replays" is a statement about where authority lives. Kafka, a managed successor, or an object-store-backed log can each satisfy it. What would not survive is moving the authority into the index itself, because then there is nothing to rebuild from and no way back from a bad mapping.
- Differentiated freshness follows human perception, not hardware. The reason stock is fast lane and a description is slow lane is that a user catches one instantly and never notices the other. That ranking of fields by perceptibility does not change when the cluster gets ten times faster; the budgets tighten, the split remains.
- Gating on independent detectors outlives any one detector. Document count, sampled diff, offline relevance score and a live canary catch four different failures. Better detectors will appear — learned diffs, interleaving, counterfactual estimators — and slot into the same gate. The principle is that promotion requires evidence and that the gate may refuse.
- The silent failure stays silent. A change that never arrives will still produce no error in 2036. Sampling against the source, and alarming on an absence of signal rather than a threshold on a rate, remains the only way to see it, whatever the pipeline is made of.
Non-functional targets
Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and the means separately.
| Quality | Target | How it is met | View |
|---|---|---|---|
| Query availability | ≥ 99.99% monthly | Query tier provisioned for peak across three AZs, replicas serving, partial results declared partial rather than silently short | 16 |
| Ingestion accept availability | ≥ 99.95% monthly | Acknowledge on durable log commit, independent of assembly and the cluster | 13 |
| Fast-lane freshness | p95 ≤ 2 s, p99 ≤ 10 s source commit to searchable | Separate fast-lane queue and writer doing partial field updates, no join and no enrichment | 15 |
| Suppression freshness | p99 ≤ 5 s | Its own lane ahead of the fast lane, so moderation cannot queue behind a price storm | 15 |
| Slow-lane freshness | p95 ≤ 60 s, p99 ≤ 5 min | Whole-document assembly with enrichment timeouts and previous-value fallback | 10 |
| Owner read-your-writes | ≤ 1 s on the owner's own query | Short-TTL recent-writes table overlaid at the query service after retrieval | 20 |
| Query latency | p50 ≤ 40 ms, p95 ≤ 120 ms, p99 ≤ 250 ms in-region | Pre-retrieval filtering, cursor pagination, per-tenant cost ceilings, reserved capacity away from reindex | 20 |
| Ingestion latency | p99 ≤ 50 ms single change, ≤ 200 ms per 500-change batch | Schema check, version guard against assembly state, durable append, acknowledge | 13 |
| Change throughput | 25,000/s steady, 100,000/s for 10 min (4x) | Partitioned log and horizontally scaled assembly with per-entity ordering only | 07 |
| Query throughput | 15,000/s steady, 45,000/s at peak (3x) | Provisioned rather than autoscaled, because scaling latency is visible inside a 250 ms p99 | 16 |
| Fan-out completion | 2,000,000 children ≤ 30 min at p95 | Rate-limited resumable job with progress and ETA, bounded synchronous enqueue | 15 |
| Full rebuild | 38,000,000 documents ≤ 4 h at ≥ 8,000 docs/s | Snapshot replay plus log catch-up, per-partition checkpoints, reserved reindex capacity share | 14 |
| Rebuild impact on queries | ≤ 15% query p99 rise for the rebuild's duration | Throttled reindex writer against a reserved query capacity share | 14 |
| Rollback | RTO 1 min | Previous index version retained 72 h; rollback is an alias move | 14 |
| Regional failover | RTO 15 min, stale but declared | Follower index in a read region, no local writes, indexing resumes from the log | 16 |
| Durability | RPO 5 s for accepted changes, RPO 0 for definitions and aliases | Replicated log; registry with point-in-time recovery; indices deliberately not backed up | 11 |
| Correctness | ≤ 0.01% documents divergent at any time | Continuous reconciliation sampling 1,000,000 documents/day, divergence published as a rate | 17 |
| Honesty | 100% of responses carry index version and observed freshness | Stamped at the query service from the resolved alias and the index's measured lag | 20 |
| Unit cost | ≤ $0.35 per M change events, ≤ $0.90 per M queries, ≤ $0.30 per GB-month | Cost attribution per tenant and index, dry run before promotion, query-to-write ratio published | 17 |
Scope
In scope
- Index definition as immutable versioned configuration — document shape, sources and join keys, per-field analysis, freshness classes, retention, replicas and relevance configuration — with compatible-versus-breaking change classification and a dry run
- Change capture from source change streams and a signed push API, with dedup, version-guarded apply, per-entity ordering and poison quarantine
- A durable partitioned change log as the platform's own system of record for what changed, with 7 days hot and 90 days archived
- Document assembly: declared-key joins, bounded resumable parent fan-out, enrichment with timeout and previous-value fallback, deterministic output
- Differentiated freshness lanes with partial document update, and suppression as its own lane
- Full reindex from snapshot plus log with per-partition checkpointing, catch-up before swap, a four-detector acceptance gate, alias swap and alias rollback
- Relevance configuration, judgement sets, offline evaluation and two live variants behind one alias
- A query API with mandatory pre-retrieval filtering, cursor pagination, cost ceilings, owner-write overlay, and index version plus observed freshness on every response
- Multi-tenancy by index, credential, quota and telemetry, with a declared shared-versus-dedicated placement policy
- Continuous sampled reconciliation, quiet-period alarms per source binding, and cost attribution per tenant and index
Explicitly out of scope
- The product's own search UI, ranking surfaces and merchandising rules
- Semantic retrieval, embeddings and vector indexes — an adjacent use case with its own index contract
- The systems of record themselves, and any authoritative copy of a business entity
- The analytics warehouse, and reporting over the engagement log beyond relevance evaluation
- A learned ranking model, its training pipeline and its feature store
- Crawling or acquiring content the platform's tenants do not already own
What a four-week prototype should prove
The prototype's job is to falsify the three central claims — that an alias swap really is a free rollback under live traffic, that two freshness lanes can be independent in practice rather than on a diagram, and that a gate of mechanical detectors catches a bad index without a human — on one tenant and a few million documents, not to build a platform.
- One tenant, one index, 3,000,000 documents from two sources: one with a real change stream, one pushing changes without a source version, so both ingestion contracts are exercised
- Fast and slow lanes on separate queues and writers, with a measured end-to-end freshness distribution for each
- A full rebuild into a second index with per-partition checkpoints, the four-detector gate, and an alias swap and rollback
- Continuous reconciliation against both sources, publishing a divergence rate
- Parent fan-out as a tracked rate-limited job over a 500,000-child parent
- The owner-write overlay on authenticated queries, with a cross-principal leak test
- An assembly worker is killed repeatedly mid-rebuild under load: every resume continues from its checkpoint, and total rebuild time degrades by less than the checkpoint interval per kill rather than restarting
- A 4x burst is applied to the slow lane alone while the fast lane runs at steady state: the fast lane's source-commit-to-searchable p99 does not move, which is the whole claim of the two-lane split
- A candidate index is built with one source partition deliberately stalled, a second with a changed analyser, and a third with a deliberately degraded boost set: the gate must refuse all three and name a different detector for each
- An alias is swapped and rolled back under live read traffic while cursors are outstanding: cursors issued against the previous version are rejected rather than silently answered from a different index
- 1% of change events are dropped at the capture tier for a week: the divergence rate rises visibly above its baseline, and the time-to-detection is measured against the chosen sample rate rather than assumed
- A brand rename touching 500,000 children runs while a merchant edits a price on one of them: the fan-out job does not revert the fast-lane field, and the field-level version guard is what prevents it
- The registry is made unavailable while queries continue: the query path serves from its last-known-good alias and relevance cache and fails static, and the measured propagation ceiling becomes the published rollback time
Open risks, carried rather than hidden
| Risk | If it lands | Response |
|---|---|---|
| The log cannot replay fast enough to rebuild the largest index inside the RTO | The central boundary becomes decorative: rebuild-and-swap stops being routine, in-place repair returns, and the rollback story goes with it | Treat replay throughput as a tested capability with a rehearsal on a schedule, and size snapshot frequency — not log retention alone — against the rebuild budget |
| Judgement sets are not maintained | The relevance detector in the swap gate becomes a rubber stamp that is harder to argue with than no gate at all, and relevance regressions ship with evidence attached | Make judgement-set freshness a published signal per index, curate from sampled queries, and let the gate report "no usable judgement set" rather than a passing score |
| A source cannot produce a comparable digest or sequence | Lost-document detection rests entirely on sampling, so the divergence bound becomes a function of sample rate and the rebuild cadence has to carry the rest of the risk | Negotiate a sequence or digest at integration time as a binding requirement, and where it is refused, raise the rebuild cadence for that index and say so in its tenant-facing contract |
| Two lanes writing the same document collide | A partial update lands on a document the slow lane is about to replace, and a field silently reverts — a divergence that reconciliation finds late and users find first | Field-level version guards on every write, with a reconciliation assertion specifically for fast-lane fields on slow-lane-written documents |
| The registry becomes a query-path dependency | Alias and relevance resolution per request makes a registry outage a query outage, and its cache TTL silently bounds how fast a rollback propagates | Serve from a last-known-good cache that fails static rather than closed, publish the propagation ceiling as the real rollback time, and keep the registry's own availability target above the query path's |
| Dedicated-cluster demand arrives before the placement policy does | Tenants are migrated ad hoc, isolation becomes a per-tenant negotiation, and the cost of unused headroom in many small clusters appears without the compliance benefit being claimed | Write the placement policy before the second large tenant onboards, and make live migration an exercised path rather than a documented one |
The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 5 areas, each with the alternatives that lost and what the choice costs.