AI Agent Orchestration Platform · View 12 of 32 · 3 · Data
The organising idea
- Storage is chosen per job rather than forced into one database, and each store is placed in one of three zones by what its loss would cost
- Only the authoritative zone needs disaster recovery; the rebuildable zone is restored by reprocessing, which is what makes the RTO achievable
- Large payloads never sit in the transactional stores — the execution record holds a reference, not a blob
Targets
- RPO 5 minutes and RTO 15 minutes for execution state; RPO effectively zero for definitions via zone-redundant SQL with a geo-replica
- Traces 90 days hot and 2 years archived; audit 7 years write-once; prompts and completions 30 days when capture is enabled
- Index rebuild from source measured at under 4 hours for the full knowledge corpus
Risks
- Rebuild time is only credible if it is exercised; a quarterly restore drill is a delivery commitment, not an aspiration
- Retention is policy-driven per tenant, so a misconfigured policy is a compliance event — retention changes require approval and are audited