AI Executive Office — CXO Assistant Platform (On-Premises)  ·  View 21 of 30  ·  6 · Operations

Deployment Topology

Where it runs, what the failure domains are, and why a second data centre is a tier option rather than a default.

Editable source SVG draw.io All views
Primary data centre — in country Application tier — spread over 3 fault domains HAProxy cluster VRRP · active pair Kong Gateway 3 nodes Kubernetes 3–30 pods Temporal 3 nodes · 3 domains AI tier vLLM on GPU nodes 8 × H100 · MIG OpenSearch 3 data · 3 shards KServe endpoints 2 nodes minimum Data tier PostgreSQL Patroni · 3 nodes Trino + Spark MinIO erasure coded Neo4j 3 nodes Redis cluster 3 nodes Evidence store object lock Hub network Egress firewall egress control Internal DNS split horizon Bastion no public admin MPLS link to the enterprise Secondary site — only where the tenant's tier permits it PostgreSQL standby RPO 5 min Evidence replica second site Redeploy from Git RTO 4 h · cold Sovereign tenants: no second site RTO is a restore, in country Enterprise network MPLS link Mail and calendar on-site servers GPU supply is the binding constraint size and order before contract private log shipping Deployment Topology — Sites and Failure Domains Interface / broker Application we own Data store Security / platform Risk / gap External / third party synchronous event / async Availability is bought inside one site, across racks and power feeds. A second site is a tier option, because for a sovereign tenant a second building may be a contract question rather than a resilience feature. v 1.0 · owner Data & AI Global Practice · date 2026-09

Decisions

  • Availability is bought inside one site, by spreading each tier across racks on independent power and network paths. For a sovereign tenant, a second building may be a contract question rather than a resilience feature, so cross-site failover is opt-in per tenant
  • Where there is no second site, recovery is a restore in the same building: the cluster rebuilt from Git by Argo CD onto spare or replacement hardware, data restored from on-site backups. RTO is measured in hours, depends on a hardware spares contract, and is stated in the contract rather than implied
  • No public endpoint on any data or AI service. Private endpoints with private DNS throughout, administrative access through Bastion with just-in-time elevation

Numbers

  • Platform availability target 99.9% monthly, measured on the ability to answer a KPI question — not on resource health
  • Rack or power-feed failure: no data loss, degraded capacity. Site loss with the DR option: RPO 5 minutes, RTO 4 hours, warm standby
  • Site loss without it: RTO governed by restore time and hardware availability, target 8 hours, tested quarterly

The open risk

  • GPU supply, and the power and cooling available in the target hall, are the binding constraints on this entire design. Accelerator lead times run to months and a dense GPU rack changes the facility's power draw, so both are confirmed with the customer's facilities team before contract, not assumed from this diagram — see the decision record for the three fallbacks