Feature Store  ·  View 15 of 21  ·  Operations

Deployment Architecture

Three availability zones for serving, a warm second region at 10%, and an RTO met by rebuilding rather than failing state over.

Editable source SVG draw.io All views
AWS eu-west-1 — active Ingress and identity Internal NLB VPC only Catalogue ALB WAF + OIDC IAM / IRSA KMS customer keys AZ-a Serving pods EKS, gRPC Flink task managers Cache node AZ-b Serving pods EKS, gRPC Flink task managers Cache node AZ-c Serving pods EKS, gRPC Flink job manager Aurora writer Regional services DynamoDB on-demand S3 + Iceberg Kinesis 7-day retention EMR Serverless 2 pools Firehose AWS eu-central-1 — standby Serving pods warm at 10% DynamoDB global table replica S3 replica cross-region Model services same VPC GitHub Actions gRPC replicate Feature Store — Deployment Architecture Interface / broker Security / platform Application we own Data store Queue / topic External / third party synchronous event / async Three AZs for serving, a warm second region at 10% capacity. S3 replicates cross-region on the same path as DynamoDB; RTO 15 min is met by rebuilding the online store, not by failing state over. v 1.0 · owner Data Platform Architecture · date 2026-09

Decisions

  • Loss of one AZ has no SLO impact; loss of the region is met by rebuilding the online store in the second region from replicated offline state, because the online store is derived and cheap to recreate.
  • The second region runs warm at 10% capacity rather than cold or hot — cold cannot meet a 15-minute RTO, hot doubles the most expensive tier in the platform.
  • Two EMR Serverless pools: one for scheduled materialisation, one for backfill. Sharing them is how a backfill takes the morning batch wave with it.

Assumptions

  • RTO 15 min for serving in the second region; registry RTO 1 h; full online rebuild ≤ 90 min.
  • DynamoDB on-demand capacity absorbs the stated 3× burst for 10 minutes without pre-provisioning.

Risks

  • A 15-minute RTO that depends on a rebuild is only as good as the last time the rebuild was run in anger.
  • Warm-at-10% means the first minutes after a regional failover are served by a tier that is scaling, so the latency SLO is breached before it is met.