| Change log — the platform's system of record for what changed |
Amazon MSK, partitioned by hash of (source, entity key), 7 days hot with tiered archive to S3 for 90 days |
Amazon Web Services |
Confluent Cloud or self-managed Kafka; Azure Event Hubs with Capture; Google Pub/Sub with a replay subscription |
Per-entity ordering without global ordering, retention long enough to cover the gap since the daily snapshot, and replay fast enough to rebuild the largest index inside the RTO |
ADR-02 |
| Change capture from relational and key-value sources |
AWS DMS for logical-decoding capture from Aurora PostgreSQL; DynamoDB Streams for the inventory source; API Gateway plus Lambda for the signed push path |
Amazon Web Services |
Debezium on Kubernetes; Azure Data Factory change feeds; Datastream on Google Cloud |
Log-based capture rather than polling, because polling cannot see deletes and cannot be asked of a production database at this size |
ADR-02 |
| Search indices, aliases and query execution |
Amazon OpenSearch Service, one domain across three AZs, 220 GB primary with two replicas, sized to hold a second copy of the largest index |
Amazon Web Services |
Elastic Cloud; self-managed OpenSearch on EKS; Azure AI Search; Vespa for a learned-ranking future |
Native atomic alias actions, partial document update, and per-index resource controls — the three engine features ADR-01, ADR-04 and ADR-16 all depend on |
ADR-01 |
| Index definition registry, aliases, promotion history, judgement sets |
Aurora PostgreSQL, multi-AZ with point-in-time recovery |
Amazon Web Services |
Cloud SQL or Spanner; Azure SQL; CockroachDB self-hosted |
The smallest and most precious store in the package needs strong consistency, RPO 0 and real relational constraints over definition versions and alias state |
ADR-11 |
| Assembly state — applied versions, join partners, parent mapping, enrichment cache |
DynamoDB with conditional writes as the version guard, plus a 60-second-TTL recent-writes table for the owner overlay |
Amazon Web Services |
Bigtable; Azure Cosmos DB; Redis with persistence for the overlay only |
A conditional write per (entity, source) is the version guard, and it must be enforced at the store rather than in process memory |
ADR-09 |
| Assembly, writers, query service, reconciler |
ECS Fargate services, scaled independently per lane, with reindex capacity rented for the duration of a rebuild |
Amazon Web Services |
EKS; Cloud Run; Azure Container Apps |
Lane independence requires independent scaling, and a four-hour rebuild should rent capacity rather than own it |
ADR-04 |
| Reindex orchestration with checkpoints and the acceptance gate |
Step Functions driving snapshot replay, log catch-up, the four detectors and the alias action, with per-partition checkpoints in DynamoDB |
Amazon Web Services |
Temporal; Argo Workflows; Azure Durable Functions |
A multi-hour, resumable, auditable process whose state must survive task replacement, and whose gate step must be able to refuse |
ADR-10 |
| Snapshot store, archive, engagement log, quarantine, audit |
S3 for daily source snapshots and retired-index archives; Kinesis Data Firehose to S3 with Athena for the query and engagement log; S3 Object Lock for the audit store |
Amazon Web Services |
GCS plus BigQuery; Azure Blob Storage with immutability policies plus Synapse |
Immutability for audit, cheap retention for evidence, and query-in-place for relevance evaluation without standing infrastructure |
ADR-14 |
| Identity, keys and the query-path perimeter |
IAM roles per workload, KMS with tenant-scoped keys where obligations require, API Gateway for authn and quotas, WAF and CloudFront at the edge |
Amazon Web Services |
Workload identity on Kubernetes with Vault; Entra ID with Azure Front Door; Google Cloud Armor with Apigee |
Four separable privileges — query, write a source, change a definition, move an alias — need four separable grants with per-workload identity |
ADR-14 |
| Observability, freshness and cost signals |
CloudWatch metrics and alarms with Managed Prometheus and Grafana for the end-to-end freshness and divergence dashboards; Cost and Usage Report tagged per tenant and index |
Amazon Web Services |
Self-hosted Prometheus, Mimir and Grafana; Azure Monitor; Google Cloud Monitoring |
End-to-end per-lane freshness and a divergence rate are derived signals, not engine metrics, and the cost attribution has to reach per-index granularity |
ADR-12 |