CI/CD Platform  ·  View 16 of 22  ·  Operations

Deployment Architecture

Three AZs of capacity, one region of evidence, and a recovery region that holds proof rather than running jobs.

Editable source SVG draw.io All views
AWS eu-west-1 — primary Control plane — EKS across three AZs Admission + API pods stateless Scheduler pods leader per shard Gate service pods Attestor pods Execution plane — AZ-a Isolation hosts on-demand Interruptible hosts background tier Execution plane — AZ-b and AZ-c Isolation hosts Interruptible hosts GPU and arm64 pools Regional data Aurora, multi-AZ Artefact + log buckets Key store AWS eu-central-1 — recovery Warm standby Control plane, scaled to zero Replicated artefacts + log Aurora read replica Source control (SaaS) Dependency mirror cross-region dispatch CI/CD Platform — Deployment Architecture Application we own Decision point Security / platform Data store External / third party event / async synchronous The recovery region holds evidence and a scaled-to-zero control plane. In-flight jobs are not failed over; they are re-dispatched. v 1.0 · owner Platform Engineering · date 2026-09

Decisions

  • The recovery region holds replicated artefacts, attestations and a read replica, with a control plane scaled to zero. It exists to serve new runs and deployments after a regional loss, not to fail over in-flight jobs.
  • In-flight jobs at the moment of a regional loss are re-dispatched, never reported successful. Losing work is acceptable; losing the truth about work is not.
  • Interruptible capacity is confined to the background tier, so a reclamation never lengthens a run a human is waiting on.

Assumptions

  • 40,000 concurrent job slots at peak; 3,200 jobs/minute sustained for 15 min with 4× burst tolerated for 5 min.
  • Regional loss: new runs and deployments within RTO 30 min, full service ≤ 2 h.
  • ≥ 55% of eligible background job-minutes on interruptible capacity; warm-pool idle waste ≤ 8% of total compute minutes.

Risks

  • Nested-virtualisation-capable instance families are a narrower market than general compute. A capacity shortage in one family is a platform-wide throughput event, which argues for two families qualified at all times.