| Job isolation |
Nested-virtualisation-capable EC2 families running a lightweight microVM hypervisor, one VM per job |
AWS + open source |
gVisor or hardened containers on GKE; Azure Container Instances with Hyper-V isolation |
Only hardware virtualisation answers "a build rooted the kernel", and this is the substrate where per-job microVMs are cheap enough for 1.9M jobs a day |
ADR-02 |
| Warm pool |
Per-host pool of pre-booted, tenant-agnostic microVMs sized from arrival rate |
Built |
Per-tenant pre-warmed pools with tenant image layers |
One fungible pool keeps the idle-waste ceiling a single manageable number across 610 tenants with uneven arrival |
ADR-10 |
| Job queue |
Amazon SQS with a durable accept-before-acknowledge contract |
AWS managed |
Kafka or NATS with consumer groups; Cloud Tasks |
An accepted run must survive any single component's loss, and the queue is work-to-do rather than the account of what happened |
ADR-12 |
| Run metadata |
Aurora PostgreSQL, multi-AZ, RPO 0 / RTO 15 min |
AWS managed |
Spanner; Azure SQL with a failover group |
Dispatch decisions are transactional — lease, entitlement, state — and need strong consistency on the hot path |
ADR-12 |
| Run event log |
Amazon Kinesis, ordered per run, retained 400 days |
AWS managed |
Kafka with a compacted history topic; Pub/Sub with an archive sink |
History, replay and recovery need an ordered immutable account that survives the loss of current state |
ADR-12 |
| Artefact store |
S3, content-addressed by digest, retention class written at seal time |
AWS managed |
GCS with object holds; ADLS Gen2 with immutability policies |
Immutable bits with a mutable pointer is what makes promote-by-digest and rollback-by-pointer possible |
ADR-15 |
| Transparency log and audit |
S3 with Object Lock in a separate custody account, 7 years, externally verifiable |
AWS managed |
A Sigstore-style public log; an append-only ledger service |
Evidence held only by the platform it exonerates is not evidence; write-once and externally checkable is the whole requirement |
ADR-08 |
| Signing |
Hardware-backed KMS key reachable only by the attestor's role |
AWS managed |
CloudHSM for offline custody; an in-cluster signer with SPIRE identity |
The signing key must be unreachable from any isolation host, which is a network and IAM property rather than a cryptographic one |
ADR-06 |
| Job identity |
OIDC workload identity minted per job, exchanged with STS for short-lived roles |
AWS + OIDC |
SPIFFE/SPIRE with a workload API; Azure workload identity federation |
A leaked credential's useful life should be the job's life, and its authority the job's authority |
ADR-07 |
| Control plane runtime |
EKS across three AZs, scheduler sharded with a leader per shard |
AWS managed |
ECS on Fargate; GKE Autopilot |
Admission, gating and attestation must scale independently, and the scheduler needs cross-tenant state that shards rather than replicates |
ADR-09 |
| Cache and mirror |
S3-backed content-keyed cache, read-wide and write-privileged, plus a regional dependency mirror |
AWS + built |
A registry-backed cache; a self-hosted Artifactory or Nexus mirror |
Splitting the write privilege removes the most-exploited CI vulnerability class, and the mirror is what makes restrictive egress fast enough to keep |
ADR-04 |
| Egress control |
No default route from sandboxes; a policy-enforcing proxy with per-tenant allow-lists |
Built |
Cloud-native egress firewall rules; a service mesh egress gateway |
The restrictive posture has to be the default one, and denial has to read as a policy decision rather than a network fault |
ADR-05 |
| Gate decisions |
Policy decision service in the control plane, versioned bundles cached per tenant |
Built + open-source policy engine |
In-pipeline gate jobs; admission policy in the target platform |
A pipeline author must not be able to weaken the gates applying to their own deployment |
ADR-16 |
| Interruptible capacity |
Spot instances for the background tier only, drain notice honoured, re-dispatch to on-demand |
AWS managed |
Spot for everything with transparent retry; on-demand only |
Cheap capacity is paid for in tail latency, so it belongs where nobody is waiting at the end of the job |
ADR-11 |
| Logs and metrics |
Agent-streamed logs tiered from searchable to archival on S3, class-based retention |
AWS + built |
OpenSearch for the whole window; a managed log product at full retention |
6 TB a day is the fastest-growing cost line, and retention by pipeline class is the only way to spend it where it is wanted |
ADR-17 |