AI Executive Office — CXO Assistant Platform (On-Premises) · View 21 of 30 · 6 · Operations
Decisions
- Availability is bought inside one site, by spreading each tier across racks on independent power and network paths. For a sovereign tenant, a second building may be a contract question rather than a resilience feature, so cross-site failover is opt-in per tenant
- Where there is no second site, recovery is a restore in the same building: the cluster rebuilt from Git by Argo CD onto spare or replacement hardware, data restored from on-site backups. RTO is measured in hours, depends on a hardware spares contract, and is stated in the contract rather than implied
- No public endpoint on any data or AI service. Private endpoints with private DNS throughout, administrative access through Bastion with just-in-time elevation
Numbers
- Platform availability target 99.9% monthly, measured on the ability to answer a KPI question — not on resource health
- Rack or power-feed failure: no data loss, degraded capacity. Site loss with the DR option: RPO 5 minutes, RTO 4 hours, warm standby
- Site loss without it: RTO governed by restore time and hardware availability, target 8 hours, tested quarterly
The open risk
- GPU supply, and the power and cooling available in the target hall, are the binding constraints on this entire design. Accelerator lead times run to months and a dense GPU rack changes the facility's power draw, so both are confirmed with the customer's facilities team before contract, not assumed from this diagram — see the decision record for the three fallbacks