CI/CD Platform · View 16 of 22 · Operations
Decisions
- The recovery region holds replicated artefacts, attestations and a read replica, with a control plane scaled to zero. It exists to serve new runs and deployments after a regional loss, not to fail over in-flight jobs.
- In-flight jobs at the moment of a regional loss are re-dispatched, never reported successful. Losing work is acceptable; losing the truth about work is not.
- Interruptible capacity is confined to the background tier, so a reclamation never lengthens a run a human is waiting on.
Assumptions
- 40,000 concurrent job slots at peak; 3,200 jobs/minute sustained for 15 min with 4× burst tolerated for 5 min.
- Regional loss: new runs and deployments within RTO 30 min, full service ≤ 2 h.
- ≥ 55% of eligible background job-minutes on interruptible capacity; warm-pool idle waste ≤ 8% of total compute minutes.
Risks
- Nested-virtualisation-capable instance families are a narrower market than general compute. A capacity shortage in one family is a platform-wide throughput event, which argues for two families qualified at all times.