AI Agent Orchestration Platform · View 23 of 32 · 5 · Operations
Decisions
- Control plane is active in the primary region with a warm standby; execution is region-pinned so a run never crosses a residency boundary mid-flight
- Zone redundancy is the primary resilience mechanism; the second region exists for region loss and for residency, not for everyday load
- Failover replays from the last checkpoint in the surviving region — it does not migrate a live run, because a live run holds provider affinity
Targets
- Control plane 99.9 percent; RTO 15 minutes and RPO 5 minutes for execution state
- Front Door health probe interval 30 seconds; automated origin failover, manual data-tier failover with a documented runbook
- Quarterly disaster-recovery exercise with a measured, published restore time
Risks
- Cosmos DB is single-write-region for consistency; a write-region loss means a manual failover inside the RTO, not an automatic one
- GPU capacity for self-hosted models is region-constrained; the standby region carries no GPU pool and those routes degrade to Foundry
- Warm standby costs roughly 18 percent of primary; a cold standby would halve that and miss the RTO