Customer 360 & Real-Time Risk Intelligence Platform  ·  View 15 of 20

Deployment Topology

What runs where across three availability zones and a warm standby region, and what survives the loss of each failure domain.

Editable source SVG draw.io All views
Primary Region — eu-west-1
Primary Region — eu-west-1
Availability Zone A
Availability Zone A
Kafka Brokers
12 nodes
Kafka Brokers...
Flink Task Managers
Autoscaled
Flink Task Managers...
Spark Workers
Spot + on demand
Spark Workers...
Availability Zone B
Availability Zone B
Kafka Brokers
12 nodes
Kafka Brokers...
Flink Task Managers
Flink Task Managers
Spark Workers
Spark Workers
Availability Zone C
Availability Zone C
Kafka Brokers
12 nodes
Kafka Brokers...
Airflow Scheduler
Active / standby
Airflow Scheduler...
Profile API Pods
Profile API Pods
Regional Managed Services
Regional Managed Services
Storage & catalog
Storage & catalog
Object Storage
Cross region replication
Object Storage...
Catalog Metastore
Multi AZ
Catalog Metastore...
State & serving
State & serving
Profile Key Value Store
3 AZ replicas
Profile Key Value Store...
SQL Warehouse
Serverless pools
SQL Warehouse...
DR Region — eu-central-1 (warm standby)
DR Region — eu-central-1 (warm standby)
Standby compute
Standby compute
Kafka Mirror Cluster
Cluster Linking
Kafka Mirror Cluster...
Pre-provisioned Compute
Scaled to zero
Pre-provisioned Compute...
Standby data
Standby data
Replicated Buckets
RPO 15 min
Replicated Buckets...
Catalog Replica
Catalog Replica
Global Traffic Manager
Health based failover
Global Traffic Manager...
async mirror
async mirror
CRR, 15 min
CRR, 15 min
metadata sync
metadata sync
primary
primary
on failover
on failover
Deployment Topology and Failure Domains
Deployment Topology and Failure Domains
Queue / topic
Queue / topic
Application we own
Application we own
Data store
Data store
Security / platform
Security / platform
Interface / broker
Interface / broker
event / async
event / async
synchronous
synchronous
failure / alternate
failure / alternate
Loss of one AZ costs capacity, not availability. RTO 4 h, RPO 15 min for a region loss.
Loss of one AZ costs capacity, not availability. RTO 4 h, RPO 15 min for a region loss.
v 1.0 · owner Platform Engineering · date 2026-08
v 1.0 · owner Platform Engineering · date 2026-08
Text is not SVG - cannot display

Availability targets

  • 99.95% for the streaming path; 99.9% for the SQL warehouse
  • Loss of one availability zone costs capacity, not availability
  • Kafka replication factor 3 with minimum in-sync replicas of 2

Disaster recovery

  • Warm standby region: RTO 4 hours, RPO 15 minutes for a full region loss
  • Cluster linking mirrors topics; object storage replicates cross-region
  • DR is exercised quarterly with a real failover, not a tabletop review

Cost posture

  • Spot capacity for batch and backfill, on-demand floors for SLA-bound jobs
  • Standby region compute is pre-provisioned but scaled to zero until needed
  • Serverless warehouse pools auto-suspend; storage tiering after 90 days