Architecture One-Pager
Solution Architecture v1.0 · Amazon Web Services · Platform Architecture · 2026-09
CI/CD Platform · Solution Architecture v1.0 · Amazon Web Services · Platform Architecture · 2026-09
Tenant code executes only inside single-use, hardware-isolated sandboxes that hold no long-lived credential and cannot write shared state; every trusted act — brokering a secret, signing provenance, deciding a gate, recording an outcome — happens in a control plane that never executes tenant code.
An engineer pushes a commit, switches to Slack, and ninety seconds later a green check appears next to their pull request. That check is the most-used piece of software in their working day and the one they think about least. Behind it, a stranger's code — because a pull request from a fork is a stranger's code — has just been compiled, tested and thrown away on the company's own machines, and a signed statement now exists saying which commit produced which container image. The problem this platform solves is not making builds fast. It is that the convenient way to build this system and the safe way differ by one boundary, and almost every widely-used CI system has shipped a variant of the same bug: tenant code was allowed to touch something the platform needed to trust. A fork build reads the instance metadata endpoint and leaves with a deployment credential. A cache entry written by that fork build is restored, an hour later, into the build that produces the production image. The second problem is downstream of the first: without a non-forgeable link between commit, build and deployed artefact, the question "what is running in production and which source produced it" is answered by archaeology rather than by query.
Classify every run at admission as trusted or untrusted from the event itself, and carry that classification, immutable, for the run's whole life. Execute every job in a sandbox created for that job and destroyed after it, with hardware-level isolation between concurrent jobs and no long-lived credential inside any of them. Keep the cache readable by everything and writable only by trusted runs. Put the signing key, the secret broker, the gate decision service and the run state machine in a control plane that never executes tenant code, so provenance records what the platform observed rather than what a build claimed. Seal each artefact by content digest with a signed attestation in an append-only log. Then promote that same digest through environments, binding configuration at deployment, so what reached production is byte-identical to what passed staging, and rollback is a pointer move rather than a rebuild.
What it is, and what it is not
- One sandbox per job, destroyed after it — not A pool of reused containers with a cleanup step
- Provenance signed by the control plane from what it observed — not A build that signs a statement about itself
- A trust class assigned once, at admission, from the event — not A permission the pipeline can declare or escalate
- A cache read by everything and written by trusted runs only — not A shared cache that everyone contributes to
- One artefact digest promoted through every environment — not A rebuild per environment from the same commit
- Gates decided outside the pipeline that they govern — not Gate jobs inside the pipeline the author controls
- Platform failure and user failure as distinct outcomes — not A red build whose cause the engineer has to guess
The decisions that are the architecture
- Every sandbox is single-use, and that is a security decision rather than a hygiene one (ADR-01) — A sandbox is created for one job and destroyed after it. Reuse is the cheapest performance win available in a CI platform and the costliest security decision available, because a residue is not a bug you find in testing. The platform buys cold-start engineering and a warm pool instead, and accepts the cost-per-job-minute that follows.
- Isolation is hardware-level, and it does not vary by trust class (ADR-02) — Jobs run in microVMs on nested-virtualisation-capable hosts, not in hardened containers, because only virtualisation answers "a build rooted the kernel". A trusted branch build is contained exactly as strictly as a fork build, because a compromised dependency inside a trusted build is indistinguishable from a hostile fork — and that is now the more likely attack.
- Trust is a property of the run, assigned at admission, derived from the event (ADR-03) — The classification is computed once from the trigger, before any capacity is committed, stored as a column on the run, and cannot be raised by anything the pipeline declares or by anything that happens later. A trust class that can be escalated mid-run is not a boundary.
- The cache is read wide and written narrow (ADR-04) — Any run may read the cache; only a trusted run may write it, and every entry is integrity-verified on restore with a failure treated as a miss. A cache writable by an untrusted run is a code-execution path into every trusted build that reads it afterwards, and the measurably slower fork build is the accepted price.
- Provenance asserts what the platform observed, never what the build claimed (ADR-06) — The signing key lives in the control plane and is never present inside a sandbox. The job reports its outcome; the attestor signs from the inputs the platform recorded. A compromised build can therefore produce a bad artefact but not a credible claim about one, which is the difference between an attestation worth verifying and a decoration.
- Credentials are minted per job and revoked when the job ends (ADR-07) — Every credential a job holds is short-lived, scoped to that job's tenant, repository, ref and trust class, and revoked at job end rather than left to expire. Identity federation therefore sits on the hot path of every job instead of being a configuration detail, and a leaked token's useful life is the job's life.
- Fairness is decided in the scheduler, not asked for politely (ADR-09) — Per-tenant concurrency entitlements with weighted fair queuing and lending of idle capacity, so a 900-job monorepo pipeline cannot push a five-job pipeline behind it. Queue wait inside an entitlement is an SLO; a tenant over its entitlement has no wait guarantee. Without this, the platform's fairness story is an email asking large teams to be considerate.
- The runner's report is evidence, not truth; the lease is truth (ADR-13) — Runner loss is detected by lease expiry rather than by silence, and a job is never reported successful on the basis of an absent runner. A network partition therefore resolves to re-dispatch, never to a false green, and retry happens only for platform failure with the cause recorded.
- Build once, promote the bits (ADR-15) — One artefact digest crosses every environment with configuration bound at deployment time. That is the only design in which "did we ship what we tested" has a cryptographic answer rather than a procedural one, and it is what makes rollback a five-minute pointer move instead of a twenty-minute rebuild.
- Gates are decided outside the pipeline they govern (ADR-16) — A decision service consulted by the deployer, not gate jobs inside the pipeline, because a pipeline author must not be able to weaken the gates that apply to their own deployment. Production fails closed when the service is unavailable, overrides are recorded with actor, reason and expiry, and their usage count is a watched signal.
Why this should still be right in ten years
Hypervisors, queue technologies, signing formats and the registries this platform talks to will all be replaced inside its life. These are the properties that should outlast them.
- The boundary is not a technology. "The control plane never executes tenant code" is a statement about where code runs, not about which hypervisor runs it. Firecracker, gVisor, a successor nobody has built yet, or dedicated hardware can each satisfy it, and every decision downstream — where the signing key lives, why the runner's report is validated, why the cache splits into two privileges — survives the substitution unchanged.
- Attacks move towards the supply chain, not away from it. The industry trend through the 2020s was steadily away from "the attacker compromises your server" and towards "the attacker compromises something your build trusts". Every year that trend continues makes the isolation-plus-attestation posture more obviously correct and the reused-container posture less defensible. The design is aligned with the direction of travel rather than with the current threat list.
- Content addressing outlives the store. An artefact identified by the digest of its bytes is identifiable in any store, by any tool, for as long as the hash function holds. Tags, environments and registries are pointers into that space and are cheap to move. A platform that identified artefacts by its own allocated ids would have to migrate identity every time it migrated storage.
- Separating platform failure from user failure is a cultural invariant. The infrastructure-failure ratio is not an AWS metric or a Kubernetes metric. It is the number that decides whether engineers read a red build or reflexively retry it, and it will matter identically on whatever substrate the platform runs in 2036. Designing the outcome taxonomy to make it measurable is a permanent asset.
- Fairness pressure only increases. Monorepos get larger, generated code gets more voluminous, and AI-assisted development raises pushes per engineer per day. Every one of those trends increases the ratio between the largest tenant's burst and the smallest tenant's pipeline, which is exactly the ratio weighted fair queuing exists to bound. A global-FIFO design degrades monotonically as the estate grows.
- The costs are the ones that get cheaper. The price of this architecture is cold-start latency and cost per job-minute on isolated compute. Both are properties of the substrate, and both have improved by roughly an order of magnitude over the last decade. The price of the alternative — an unprovable artefact in production — does not get cheaper with hardware.
Non-functional targets
Every figure below is a stated assumption for this design, chosen to be defensible and arguable rather than measured. A reviewer who changes one can follow it to the decision that depends on it.
| Quality | Target | How it is met | View |
|---|---|---|---|
| Trigger acceptance availability | ≥ 99.95% monthly | Stateless receivers across three AZs; a run is durable in the queue before the trigger is acknowledged | 16 |
| Cold start | Queued to first step p50 ≤ 8 s, p95 ≤ 25 s, p99 ≤ 60 s | Tenant-agnostic warm pool sized from observed arrival rate; sandbox claimed at dispatch | 13 |
| Queue fairness | p95 ≤ 30 s and no job over 10 min while inside entitlement | Weighted fair queuing across tenants with lending of idle entitlement | 02 |
| Throughput | 1.9M jobs/day, 3,200 jobs/min sustained, 4× burst for 5 min | Execution plane scales independently of the control plane; background tier shed first | 16 |
| Gate decision latency | p99 ≤ 500 ms | Decision service co-located with the control plane; policy bundles cached per tenant | 17 |
| Live log tail | Emitted to visible p95 ≤ 1.5 s | Agent streams to the control plane, which fans out and tiers to object storage | 10 |
| Rollback | p95 ≤ 5 min to a retained digest, ceiling 15 min | Environment pointer move, no rebuild; prior digests retained per class | 17 |
| Isolation correctness | Zero sandboxes serving two jobs; zero secrets reaching an untrusted run | Single-use microVM per job; secrets refused by trust class at the broker | 20 |
| Provenance integrity | Zero accepted attestations for builds this platform did not execute | Signing key held only by the control-plane attestor; append-only transparency log | 21 |
| Infrastructure-failure ratio | ≤ 0.5% of jobs weekly, alerting at 1.0% hourly | Six-outcome job taxonomy; retry only on platform failure with cause recorded | 22 |
| Durability | Run metadata RPO 0 / RTO 15 min; artefacts and attestations RPO 0 / RTO 1 h | Aurora multi-AZ; cross-region replication of evidence; caches carry no RPO | 11 |
| Regional recovery | New runs and deployments within 30 min, full service ≤ 2 h | Warm standby control plane scaled to zero; in-flight jobs re-dispatched, never assumed successful | 16 |
| Cost | ≤ $0.011 per job-minute at p50 utilisation; warm-pool idle waste ≤ 8% | Interruptible capacity for background tiers; warm pool sized from arrival rate | 18 |
| Operability | Definition change effective ≤ 60 s after merge; new repository onboarded ≤ 10 min | Definition read from the repository at the commit under test; no platform-team step | 02 |
Scope
In scope
- Repository-hosted pipeline definitions, DAG compilation, schema and policy admission, and an immutable effective-definition record per run
- Trust classification of every run at admission, and the credential, cache and publish postures that follow from it
- A durable queue with per-tenant entitlements, priority tiers and visible queue position
- Single-use, hardware-isolated sandbox execution with mediated egress, streaming logs and a six-outcome taxonomy
- A content-addressed artefact store, a read-wide write-privileged cache, and a dependency mirror
- Control-plane-signed provenance and SBOM per artefact in an append-only transparency log
- Environment registry, ordered promotion by digest, deployment locks, rollback, and a gate decision service outside the pipeline
- Run observability: live tail, first-failure summary, flake detection, delivery metrics and cost per run
Explicitly out of scope
- The source-control system itself — the platform consumes its events and API and writes back only a check status
- The public package and container registries a build pulls from; they are proxied and cached, not hosted
- The runtime platform that receives a deployment, and its own configuration stores
- Test authoring, test frameworks and test infrastructure inside a job
- Incident management, on-call routing and the postmortem process
- Developer laptops and any local build cache on them
- Progressive delivery, remote build execution and test impact analysis, all named and deferred to Phase 3
What a four-week prototype should prove
The prototype's job is to falsify the central claim — that a fork build can run usefully on the same substrate as a trusted build without ever touching anything the platform needs to trust — on two repositories and one environment, not to build a platform.
- One repository, two pipelines: a trusted branch build and a fork pull-request build, on one microVM host with a warm pool of ten
- Measure queued-to-first-step at p50 and p95 with a cold pool and a warm pool, and publish the cost per job-minute of both
- Implement the trust classifier and prove, by attempting it, that a fork build cannot obtain a secret, a cache write credential or a publish identity
- Implement the attestor with a signing key the sandbox provably cannot reach, and verify an attestation from outside the platform using only the transparency log
- Promote one digest through two environments with configuration bound at deployment, then roll back and measure it
- Implement the six-outcome taxonomy and the lease, and measure the infrastructure-failure ratio over 500 deliberately-disrupted jobs
- A fork build reads the instance metadata endpoint: it receives nothing usable, the attempt is logged against the run, and the maintainer sees why
- A fork build writes a poisoned cache entry: the write is refused by credential, not by validation, and no trusted build is affected
- The host is killed mid-job: the lease expires, the job is re-dispatched with attempt 2 and cause recorded, the partial log survives, and no green check appears
- The gate decision service is stopped: promotion to production is refused with a named reason, and promotion to test proceeds under fail-open policy
- A signed attestation is edited in place: external verification against the transparency log fails, and the artefact is quarantined rather than deleted
- Two promotions to the same environment are issued a second apart: the deployment lock serialises them and the audit trail shows both, in order
Open risks, carried rather than hidden
| Risk | If it lands | Response |
|---|---|---|
| Cold start and cost per job-minute make the safe path the slow path, so teams build elsewhere | Engineers run builds on laptops or in a side account, and the provenance chain quietly stops covering what ships | Treat queued-to-first-step and cost per job-minute as headline SLOs rather than consequences. Size the warm pool from arrival rate, publish both numbers per tenant, and accept idle waste up to 8% as the price of being the easy path. |
| Pressure to give fork builds secrets, because a maintainer needs an integration test to pass | The single most dangerous feature the platform could ship, and the one most likely to be requested with a good reason attached | Offer the trusted re-run at merge and a maintainer-triggered trusted run on an explicitly reviewed head, never a "run with secrets" toggle on the fork itself. Make the alternative good enough that the toggle is not asked for twice. |
| The control plane is both a bottleneck and a blast radius | Every trusted act — secrets, signing, gates, outcomes — depends on it, so its availability must exceed everything it serves | Shard the scheduler, keep admission and the gate service independently scalable, and design degradation so new admissions fail before running work does. Accepted runs survive in the durable queue; in-flight jobs complete without the control plane's help. |
| Nested-virtualisation-capable instance families are a narrow market | A capacity shortage in one family becomes a platform-wide throughput event with no software fix | Qualify two instance families at all times and keep the background tier on interruptible capacity that can be displaced, so a shortage costs nightly throughput before it costs interactive latency. |
| A confused-deputy bug in the attestor attributes one job's report to another's inputs | A signed attestation that is internally valid and factually wrong, which is worse than no attestation at all | Treat the attestor's input validation as the highest-review-value code in the platform, bind every report to the job's lease and identity, and keep the transparency log independently verifiable so an inconsistency is discoverable without the platform's cooperation. |
The reasoning behind every component and technology choice is in the Architecture Decision Record: 17 records across 6 areas, each with the alternatives that lost and what the choice costs.