| Offline store |
S3 with Apache Iceberg tables and the Glue Data Catalog |
AWS + open format |
Delta Lake on ADLS Gen2, or BigQuery with time travel |
Snapshot isolation lets a 45-minute training join read a stable view while materialisation commits, and the snapshot id becomes a pinnable input to the manifest |
ADR-06 |
| Online store |
DynamoDB, one item per entity and feature group, on-demand capacity |
AWS |
Cosmos DB, Bigtable, or self-hosted Cassandra / Redis |
Single-digit-millisecond reads at 320k/s with no capacity planning for the stated 3× burst, and cheap enough to treat as disposable |
ADR-08 |
| Stream log |
Kinesis Data Streams with 7-day retention |
AWS |
MSK / Kafka, Event Hubs, or Pub/Sub |
Retention is the replay window that every streaming recovery path in view 21 depends on; 7 days is the assumption the RPO rests on |
ADR-07 |
| Stream processing |
Managed Service for Apache Flink, checkpointed |
AWS + open engine |
Self-hosted Flink, Spark Structured Streaming, or Dataflow |
Windowed aggregates from 1 minute to 24 hours with exactly-once checkpointing, and a job graph the compiler can emit |
ADR-03 |
| Batch materialisation |
EMR Serverless, two pools |
AWS + open engine |
Databricks jobs, Glue, or Dataproc |
Separate pools give backfill its own capacity and pre-emption without a second cluster to keep warm |
ADR-14 |
| Registry |
Aurora PostgreSQL Serverless v2 |
AWS |
Azure SQL, Cloud SQL, or self-hosted PostgreSQL |
Small, strongly consistent and the only store whose loss is unrecoverable; relational because the model in view 11 is relational |
ADR-11 |
| Serving tier |
gRPC services on EKS behind an internal NLB, three AZs |
AWS + Kubernetes |
AKS, GKE, or ECS Fargate |
Stateless pods scaled independently of materialisation, with IRSA giving each caller a workload identity rather than a key |
ADR-17 |
| Hot-key cache |
ElastiCache with a short TTL, values immutable within their window only |
AWS |
Momento, Redis Enterprise, or an in-process cache |
Coalesces reads on a small number of very popular stores without becoming a second, undetectable staleness source |
ADR-10 |
| Serving log |
Kinesis Data Firehose to S3, partitioned by model and hour |
AWS |
Event Hubs Capture, or a direct async write to object storage |
Buffered, cheap, and classification-aware by prefix, which is what lets the log inherit the vector's maximum class |
ADR-16 |
| Offline access control |
Lake Formation column and row grants |
AWS |
Unity Catalog, Purview, or Ranger |
Scoped credentials for a human's offline read, so purpose and ACL are enforced before the data leaves the table |
ADR-17 |
| Definition pipeline |
GitHub with required review and GitHub Actions |
External |
GitLab CI, Azure DevOps, or CodeBuild |
Registration is a merge, so review, duplicate detection and the parity gate all happen where the author is already working |
ADR-12 |
| Observability |
Amazon Managed Prometheus and Managed Grafana, with CloudWatch for platform metrics |
AWS + open ecosystem |
Azure Monitor, Cloud Monitoring, or self-hosted Prometheus |
Freshness and lag are SLO series rather than logs, and they need to be queryable by feature group at 120-group cardinality |
ADR-13 |
| Encryption and keys |
KMS with customer-managed keys for restricted groups |
AWS |
Key Vault, Cloud KMS, or Vault |
A separate key per classification is the boundary that makes a classification more than a tag |
ADR-16 |
| Cross-region recovery |
S3 cross-region replication plus a DynamoDB global table replica |
AWS |
Paired-region storage replication with a rebuild job |
The offline replica is what the rebuild reads; the online replica shortens the rebuild rather than guaranteeing correctness |
ADR-15 |