CI/CD Platform

Architecture Views

22 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

A multi-tenant CI/CD platform for 4,200 engineers in 610 teams: 210,000 pipeline runs a day, 1.9M jobs, every one of them in a sandbox that is created for that job and destroyed after it. Read the set in seven acts. It opens with who is waiting for a green check and what happens when the code under test came from a stranger, because that is the decision the rest of the architecture is built around: the control plane never executes tenant code, and everything that must be trusted happens where tenant code cannot reach.

Context and scope

What sits inside the boundary, what the platform reads without writing, and what it deliberately does not own.

People and journeys

Who the platform is for, and what each of them gets to do. Three journeys, failing in three different places.
03 Inside the company Application engineer 4,200 in 610 teams Goal — Tell me in minutes whether my change is safe to merge, and if it is not, tell me what I broke rather than making me guess. Core journeys Get a green check Reproduce a red job Code reviewer every merge Goal — Let me trust the check beside the diff so I review the change instead of re-running the build. Core journeys Read the evidence on a PR Platform engineer 14 people Goal — Keep 40,000 slots busy without letting one monorepo starve everyone else, and know when a failure is ours. Core journeys Rebalance entitlements Triage infra-failure ratio Accountable for what ships Release owner 9 business units Goal — Ship exactly the bits we tested, and be able to take them back in five minutes without a rebuild. Core journeys Promote and roll back Security engineer policy owner Goal — Make the safe path the default path, so no team has to remember to be careful about a fork build. Core journeys Change a gate policy Answer an audit request Outside the company Outside contributor ~300 per quarter Goal — Get my fix tested and reviewed without being treated as an attacker, or left waiting with no reason given. Core journeys Land a fork contribution Auditor 2 reviews per year Goal — Show me, for one image in production, which commit produced it and who approved it. Core journeys Trace an image to a commit Machines in the cast Scheduled trigger 18k runs/night Goal — Use the capacity nobody else wants, and get out of the way when a human is waiting. Core journeys Run nightly on spare capacity Review bot lint and SAST Goal — Post findings on the diff fast enough to be read before the human reviewer arrives. Core journeys Post findings on a PR CI/CD Platform — Actors and Their Core Journeys Person or role Journey / task External / third party Security / platform Three journeys get their own map: the engineer's green check, the fork contribution, and the release promotion. They fail in different places. v 1.0 · owner Platform Engineering · date 2026-09 Actors and Their Core Journeys Nine actors, three of them machines, and the goals they would state in their own words. HTML page SVG draw.io

Structure

The parts, the trust boundary between them, and the surfaces they expose.

Data

What is stored, how long it lives, and which of it can be thrown away.

Runtime

What actually happens between a push and a signed artefact, and how the three trust classes differ.

Operations

Where it runs, how a release moves, what is watched, and how the loop closes.

Assurance

Why it is safe to run a stranger's code, and what is assumed to fail.

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

Tenant code executes only inside single-use, hardware-isolated sandboxes that hold no long-lived credential and cannot write shared state; every trusted act — brokering a secret, signing provenance, deciding a gate, recording an outcome — happens in a control plane that never executes tenant code.

An engineer pushes a commit, switches to Slack, and ninety seconds later a green check appears next to their pull request. That check is the most-used piece of software in their working day and the one they think about least. Behind it, a stranger's code — because a pull request from a fork is a stranger's code — has just been compiled, tested and thrown away on the company's own machines, and a signed statement now exists saying which commit produced which container image. The problem this platform solves is not making builds fast. It is that the convenient way to build this system and the safe way differ by one boundary, and almost every widely-used CI system has shipped a variant of the same bug: tenant code was allowed to touch something the platform needed to trust. A fork build reads the instance metadata endpoint and leaves with a deployment credential. A cache entry written by that fork build is restored, an hour later, into the build that produces the production image. The second problem is downstream of the first: without a non-forgeable link between commit, build and deployed artefact, the question "what is running in production and which source produced it" is answered by archaeology rather than by query.

Classify every run at admission as trusted or untrusted from the event itself, and carry that classification, immutable, for the run's whole life. Execute every job in a sandbox created for that job and destroyed after it, with hardware-level isolation between concurrent jobs and no long-lived credential inside any of them. Keep the cache readable by everything and writable only by trusted runs. Put the signing key, the secret broker, the gate decision service and the run state machine in a control plane that never executes tenant code, so provenance records what the platform observed rather than what a build claimed. Seal each artefact by content digest with a signed attestation in an append-only log. Then promote that same digest through environments, binding configuration at deployment, so what reached production is byte-identical to what passed staging, and rollback is a pointer move rather than a rebuild.

What it is, and what it is not

One sandbox per job, destroyed after itA pool of reused containers with a cleanup step
Provenance signed by the control plane from what it observedA build that signs a statement about itself
A trust class assigned once, at admission, from the eventA permission the pipeline can declare or escalate
A cache read by everything and written by trusted runs onlyA shared cache that everyone contributes to
One artefact digest promoted through every environmentA rebuild per environment from the same commit
Gates decided outside the pipeline that they governGate jobs inside the pipeline the author controls
Platform failure and user failure as distinct outcomesA red build whose cause the engineer has to guess

The decisions that are the architecture

01Every sandbox is single-use, and that is a security decision rather than a hygiene one

A sandbox is created for one job and destroyed after it. Reuse is the cheapest performance win available in a CI platform and the costliest security decision available, because a residue is not a bug you find in testing. The platform buys cold-start engineering and a warm pool instead, and accepts the cost-per-job-minute that follows.

ADR-01

02Isolation is hardware-level, and it does not vary by trust class

Jobs run in microVMs on nested-virtualisation-capable hosts, not in hardened containers, because only virtualisation answers "a build rooted the kernel". A trusted branch build is contained exactly as strictly as a fork build, because a compromised dependency inside a trusted build is indistinguishable from a hostile fork — and that is now the more likely attack.

ADR-02

03Trust is a property of the run, assigned at admission, derived from the event

The classification is computed once from the trigger, before any capacity is committed, stored as a column on the run, and cannot be raised by anything the pipeline declares or by anything that happens later. A trust class that can be escalated mid-run is not a boundary.

ADR-03

04The cache is read wide and written narrow

Any run may read the cache; only a trusted run may write it, and every entry is integrity-verified on restore with a failure treated as a miss. A cache writable by an untrusted run is a code-execution path into every trusted build that reads it afterwards, and the measurably slower fork build is the accepted price.

ADR-04

05Provenance asserts what the platform observed, never what the build claimed

The signing key lives in the control plane and is never present inside a sandbox. The job reports its outcome; the attestor signs from the inputs the platform recorded. A compromised build can therefore produce a bad artefact but not a credible claim about one, which is the difference between an attestation worth verifying and a decoration.

ADR-06

06Credentials are minted per job and revoked when the job ends

Every credential a job holds is short-lived, scoped to that job's tenant, repository, ref and trust class, and revoked at job end rather than left to expire. Identity federation therefore sits on the hot path of every job instead of being a configuration detail, and a leaked token's useful life is the job's life.

ADR-07

07Fairness is decided in the scheduler, not asked for politely

Per-tenant concurrency entitlements with weighted fair queuing and lending of idle capacity, so a 900-job monorepo pipeline cannot push a five-job pipeline behind it. Queue wait inside an entitlement is an SLO; a tenant over its entitlement has no wait guarantee. Without this, the platform's fairness story is an email asking large teams to be considerate.

ADR-09

08The runner's report is evidence, not truth; the lease is truth

Runner loss is detected by lease expiry rather than by silence, and a job is never reported successful on the basis of an absent runner. A network partition therefore resolves to re-dispatch, never to a false green, and retry happens only for platform failure with the cause recorded.

ADR-13

09Build once, promote the bits

One artefact digest crosses every environment with configuration bound at deployment time. That is the only design in which "did we ship what we tested" has a cryptographic answer rather than a procedural one, and it is what makes rollback a five-minute pointer move instead of a twenty-minute rebuild.

ADR-15

10Gates are decided outside the pipeline they govern

A decision service consulted by the deployer, not gate jobs inside the pipeline, because a pipeline author must not be able to weaken the gates that apply to their own deployment. Production fails closed when the service is unavailable, overrides are recorded with actor, reason and expiry, and their usage count is a watched signal.

ADR-16

Why this should still be right in ten years

Hypervisors, queue technologies, signing formats and the registries this platform talks to will all be replaced inside its life. These are the properties that should outlast them.

The boundary is not a technology

"The control plane never executes tenant code" is a statement about where code runs, not about which hypervisor runs it. Firecracker, gVisor, a successor nobody has built yet, or dedicated hardware can each satisfy it, and every decision downstream — where the signing key lives, why the runner's report is validated, why the cache splits into two privileges — survives the substitution unchanged.

Attacks move towards the supply chain, not away from it

The industry trend through the 2020s was steadily away from "the attacker compromises your server" and towards "the attacker compromises something your build trusts". Every year that trend continues makes the isolation-plus-attestation posture more obviously correct and the reused-container posture less defensible. The design is aligned with the direction of travel rather than with the current threat list.

Content addressing outlives the store

An artefact identified by the digest of its bytes is identifiable in any store, by any tool, for as long as the hash function holds. Tags, environments and registries are pointers into that space and are cheap to move. A platform that identified artefacts by its own allocated ids would have to migrate identity every time it migrated storage.

Separating platform failure from user failure is a cultural invariant

The infrastructure-failure ratio is not an AWS metric or a Kubernetes metric. It is the number that decides whether engineers read a red build or reflexively retry it, and it will matter identically on whatever substrate the platform runs in 2036. Designing the outcome taxonomy to make it measurable is a permanent asset.

Fairness pressure only increases

Monorepos get larger, generated code gets more voluminous, and AI-assisted development raises pushes per engineer per day. Every one of those trends increases the ratio between the largest tenant's burst and the smallest tenant's pipeline, which is exactly the ratio weighted fair queuing exists to bound. A global-FIFO design degrades monotonically as the estate grows.

The costs are the ones that get cheaper

The price of this architecture is cold-start latency and cost per job-minute on isolated compute. Both are properties of the substrate, and both have improved by roughly an order of magnitude over the last decade. The price of the alternative — an unprovable artefact in production — does not get cheaper with hardware.

Non-functional targets

Every figure below is a stated assumption for this design, chosen to be defensible and arguable rather than measured. A reviewer who changes one can follow it to the decision that depends on it.

QualityTargetHow it is metView
Trigger acceptance availability ≥ 99.95% monthly Stateless receivers across three AZs; a run is durable in the queue before the trigger is acknowledged 16
Cold start Queued to first step p50 ≤ 8 s, p95 ≤ 25 s, p99 ≤ 60 s Tenant-agnostic warm pool sized from observed arrival rate; sandbox claimed at dispatch 13
Queue fairness p95 ≤ 30 s and no job over 10 min while inside entitlement Weighted fair queuing across tenants with lending of idle entitlement 02
Throughput 1.9M jobs/day, 3,200 jobs/min sustained, 4× burst for 5 min Execution plane scales independently of the control plane; background tier shed first 16
Gate decision latency p99 ≤ 500 ms Decision service co-located with the control plane; policy bundles cached per tenant 17
Live log tail Emitted to visible p95 ≤ 1.5 s Agent streams to the control plane, which fans out and tiers to object storage 10
Rollback p95 ≤ 5 min to a retained digest, ceiling 15 min Environment pointer move, no rebuild; prior digests retained per class 17
Isolation correctness Zero sandboxes serving two jobs; zero secrets reaching an untrusted run Single-use microVM per job; secrets refused by trust class at the broker 20
Provenance integrity Zero accepted attestations for builds this platform did not execute Signing key held only by the control-plane attestor; append-only transparency log 21
Infrastructure-failure ratio ≤ 0.5% of jobs weekly, alerting at 1.0% hourly Six-outcome job taxonomy; retry only on platform failure with cause recorded 22
Durability Run metadata RPO 0 / RTO 15 min; artefacts and attestations RPO 0 / RTO 1 h Aurora multi-AZ; cross-region replication of evidence; caches carry no RPO 11
Regional recovery New runs and deployments within 30 min, full service ≤ 2 h Warm standby control plane scaled to zero; in-flight jobs re-dispatched, never assumed successful 16
Cost ≤ $0.011 per job-minute at p50 utilisation; warm-pool idle waste ≤ 8% Interruptible capacity for background tiers; warm pool sized from arrival rate 18
Operability Definition change effective ≤ 60 s after merge; new repository onboarded ≤ 10 min Definition read from the repository at the commit under test; no platform-team step 02

Scope

In scope

  • Repository-hosted pipeline definitions, DAG compilation, schema and policy admission, and an immutable effective-definition record per run
  • Trust classification of every run at admission, and the credential, cache and publish postures that follow from it
  • A durable queue with per-tenant entitlements, priority tiers and visible queue position
  • Single-use, hardware-isolated sandbox execution with mediated egress, streaming logs and a six-outcome taxonomy
  • A content-addressed artefact store, a read-wide write-privileged cache, and a dependency mirror
  • Control-plane-signed provenance and SBOM per artefact in an append-only transparency log
  • Environment registry, ordered promotion by digest, deployment locks, rollback, and a gate decision service outside the pipeline
  • Run observability: live tail, first-failure summary, flake detection, delivery metrics and cost per run

Explicitly out of scope

  • The source-control system itself — the platform consumes its events and API and writes back only a check status
  • The public package and container registries a build pulls from; they are proxied and cached, not hosted
  • The runtime platform that receives a deployment, and its own configuration stores
  • Test authoring, test frameworks and test infrastructure inside a job
  • Incident management, on-call routing and the postmortem process
  • Developer laptops and any local build cache on them
  • Progressive delivery, remote build execution and test impact analysis, all named and deferred to Phase 3

What a four-week prototype should prove

The prototype's job is to falsify the central claim — that a fork build can run usefully on the same substrate as a trusted build without ever touching anything the platform needs to trust — on two repositories and one environment, not to build a platform.

  1. One repository, two pipelines: a trusted branch build and a fork pull-request build, on one microVM host with a warm pool of ten
  2. Measure queued-to-first-step at p50 and p95 with a cold pool and a warm pool, and publish the cost per job-minute of both
  3. Implement the trust classifier and prove, by attempting it, that a fork build cannot obtain a secret, a cache write credential or a publish identity
  4. Implement the attestor with a signing key the sandbox provably cannot reach, and verify an attestation from outside the platform using only the transparency log
  5. Promote one digest through two environments with configuration bound at deployment, then roll back and measure it
  6. Implement the six-outcome taxonomy and the lease, and measure the infrastructure-failure ratio over 500 deliberately-disrupted jobs
  • A fork build reads the instance metadata endpoint: it receives nothing usable, the attempt is logged against the run, and the maintainer sees why
  • A fork build writes a poisoned cache entry: the write is refused by credential, not by validation, and no trusted build is affected
  • The host is killed mid-job: the lease expires, the job is re-dispatched with attempt 2 and cause recorded, the partial log survives, and no green check appears
  • The gate decision service is stopped: promotion to production is refused with a named reason, and promotion to test proceeds under fail-open policy
  • A signed attestation is edited in place: external verification against the transparency log fails, and the artefact is quarantined rather than deleted
  • Two promotions to the same environment are issued a second apart: the deployment lock serialises them and the audit trail shows both, in order

Open risks, carried rather than hidden

RiskIf it landsResponse
Cold start and cost per job-minute make the safe path the slow path, so teams build elsewhere Engineers run builds on laptops or in a side account, and the provenance chain quietly stops covering what ships Treat queued-to-first-step and cost per job-minute as headline SLOs rather than consequences. Size the warm pool from arrival rate, publish both numbers per tenant, and accept idle waste up to 8% as the price of being the easy path.
Pressure to give fork builds secrets, because a maintainer needs an integration test to pass The single most dangerous feature the platform could ship, and the one most likely to be requested with a good reason attached Offer the trusted re-run at merge and a maintainer-triggered trusted run on an explicitly reviewed head, never a "run with secrets" toggle on the fork itself. Make the alternative good enough that the toggle is not asked for twice.
The control plane is both a bottleneck and a blast radius Every trusted act — secrets, signing, gates, outcomes — depends on it, so its availability must exceed everything it serves Shard the scheduler, keep admission and the gate service independently scalable, and design degradation so new admissions fail before running work does. Accepted runs survive in the durable queue; in-flight jobs complete without the control plane's help.
Nested-virtualisation-capable instance families are a narrow market A capacity shortage in one family becomes a platform-wide throughput event with no software fix Qualify two instance families at all times and keep the background tier on interruptible capacity that can be displaced, so a shortage costs nightly throughput before it costs interactive latency.
A confused-deputy bug in the attestor attributes one job's report to another's inputs A signed attestation that is internally valid and factually wrong, which is worse than no attestation at all Treat the attestor's input validation as the highest-review-value code in the platform, bind every report to the job's lease and identity, and keep the transparency log independently verifiable so an inconsistency is discoverable without the platform's cooperation.

Architecture Decision Record

Why every component and every technology on these 22 views is what it is, and what each choice costs.

Seventeen decisions make up this architecture. Everything else across the twenty-two views is either a consequence of one of them or a detail that could be decided differently next quarter without anybody having to redraw the set. Each record carries more than the classic context / decision / consequences triple: the forcing question, how the decision is actually realised on the chosen stack, the alternatives including those that are right for a different organisation, the conditions that would flip the choice, why the choice should still hold as scale and technology change, and the transferable lesson.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a product company of 4,200 engineers in 610 teams across 38,000 repositories, running 210,000 pipeline runs and 1.9M jobs a day, peaking at 3,200 jobs a minute with 40,000 concurrent job slots, a median job of 3m 10s and a p95 of 22 minutes, 480 TB of retained artefacts and 6 TB a day of raw log output, with roughly 300 fork pull requests a quarter from outside the company. A reviewer who disagrees with a number can follow it to the decision that depends on it; that is what the numbers are for.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

Isolation — running code nobody reviewed 5

The decisions that make it ordinary rather than dangerous to execute a stranger's code on the company's machines.

ADR-01Every sandbox is single-use, created for one job and destroyed after it ADR-02Isolation is hardware-level, and it does not vary by trust class ADR-03Trust is a property of the run, assigned at admission from the event ADR-04The cache is read by everything and written only by trusted runs ADR-05All sandbox egress is mediated by a policy proxy, on by default

Trust and provenance 3

The decisions that make the link between commit, build and deployed artefact non-forgeable.

ADR-06The attestor signs provenance; the build never signs anything ADR-07Credentials are minted per job and revoked when the job ends ADR-08Attestations are recorded in an append-only log that can be verified without the platform

Fairness and capacity 3

The decisions that stop one tenant's burst becoming everyone else's queue.

ADR-09Fairness is enforced in the scheduler, by entitlement and weighted fair queuing ADR-10The warm pool is tenant-agnostic, and a claimed sandbox is never returned to it ADR-11Interruptible capacity is confined to the background tier

Run state and recovery 3

The decisions that make a partition resolve to re-dispatch rather than to a false green.

ADR-12The relational store is authoritative for scheduling; the event log is authoritative for history ADR-13Runner loss is detected by lease expiry, and a job is never successful on an absent runner ADR-14Six job outcomes, and a platform failure is never reported as the user's

Promotion and gating 2

The decisions that make "did we ship what we tested" answerable, and the gates unweakenable by the author.

ADR-15Build once, promote the digest, bind configuration at deployment ADR-16Gates are evaluated outside the pipeline they govern, and fail closed in production

Cost, retention and developer trust 1

The decisions that keep the safe path the cheap path, and a red build worth reading.

ADR-17Retention is set by artefact provenance class and pipeline class, not globally

Technology by capability

Amazon Web Services was chosen for this exercise deliberately, and for two reasons that pull in the same direction. The first is rotation: across the repository's previous use cases, Azure and self-hosted open-source stacks dominate, and AWS and Google Cloud are the two least leaned on. The second is that the topic genuinely has a cloud-shaped centre of gravity — build isolation at this volume needs nested-virtualisation-capable instances running a lightweight microVM hypervisor, which is where the real hosted CI vendors run and where the primitive originated. The requirement in ask.md stays vendor-neutral throughout; the table below is the architecture's answer, not the requirement's.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Job isolation Nested-virtualisation-capable EC2 families running a lightweight microVM hypervisor, one VM per job AWS + open source gVisor or hardened containers on GKE; Azure Container Instances with Hyper-V isolation Only hardware virtualisation answers "a build rooted the kernel", and this is the substrate where per-job microVMs are cheap enough for 1.9M jobs a day ADR-02
Warm pool Per-host pool of pre-booted, tenant-agnostic microVMs sized from arrival rate Built Per-tenant pre-warmed pools with tenant image layers One fungible pool keeps the idle-waste ceiling a single manageable number across 610 tenants with uneven arrival ADR-10
Job queue Amazon SQS with a durable accept-before-acknowledge contract AWS managed Kafka or NATS with consumer groups; Cloud Tasks An accepted run must survive any single component's loss, and the queue is work-to-do rather than the account of what happened ADR-12
Run metadata Aurora PostgreSQL, multi-AZ, RPO 0 / RTO 15 min AWS managed Spanner; Azure SQL with a failover group Dispatch decisions are transactional — lease, entitlement, state — and need strong consistency on the hot path ADR-12
Run event log Amazon Kinesis, ordered per run, retained 400 days AWS managed Kafka with a compacted history topic; Pub/Sub with an archive sink History, replay and recovery need an ordered immutable account that survives the loss of current state ADR-12
Artefact store S3, content-addressed by digest, retention class written at seal time AWS managed GCS with object holds; ADLS Gen2 with immutability policies Immutable bits with a mutable pointer is what makes promote-by-digest and rollback-by-pointer possible ADR-15
Transparency log and audit S3 with Object Lock in a separate custody account, 7 years, externally verifiable AWS managed A Sigstore-style public log; an append-only ledger service Evidence held only by the platform it exonerates is not evidence; write-once and externally checkable is the whole requirement ADR-08
Signing Hardware-backed KMS key reachable only by the attestor's role AWS managed CloudHSM for offline custody; an in-cluster signer with SPIRE identity The signing key must be unreachable from any isolation host, which is a network and IAM property rather than a cryptographic one ADR-06
Job identity OIDC workload identity minted per job, exchanged with STS for short-lived roles AWS + OIDC SPIFFE/SPIRE with a workload API; Azure workload identity federation A leaked credential's useful life should be the job's life, and its authority the job's authority ADR-07
Control plane runtime EKS across three AZs, scheduler sharded with a leader per shard AWS managed ECS on Fargate; GKE Autopilot Admission, gating and attestation must scale independently, and the scheduler needs cross-tenant state that shards rather than replicates ADR-09
Cache and mirror S3-backed content-keyed cache, read-wide and write-privileged, plus a regional dependency mirror AWS + built A registry-backed cache; a self-hosted Artifactory or Nexus mirror Splitting the write privilege removes the most-exploited CI vulnerability class, and the mirror is what makes restrictive egress fast enough to keep ADR-04
Egress control No default route from sandboxes; a policy-enforcing proxy with per-tenant allow-lists Built Cloud-native egress firewall rules; a service mesh egress gateway The restrictive posture has to be the default one, and denial has to read as a policy decision rather than a network fault ADR-05
Gate decisions Policy decision service in the control plane, versioned bundles cached per tenant Built + open-source policy engine In-pipeline gate jobs; admission policy in the target platform A pipeline author must not be able to weaken the gates applying to their own deployment ADR-16
Interruptible capacity Spot instances for the background tier only, drain notice honoured, re-dispatch to on-demand AWS managed Spot for everything with transparent retry; on-demand only Cheap capacity is paid for in tail latency, so it belongs where nobody is waiting at the end of the job ADR-11
Logs and metrics Agent-streamed logs tiered from searchable to archival on S3, class-based retention AWS + built OpenSearch for the whole window; a managed log product at full retention 6 TB a day is the fastest-growing cost line, and retention by pipeline class is the only way to spend it where it is wanted ADR-17

The decisions, and the alternatives that lost

Isolation — running code nobody reviewedThe decisions that make it ordinary rather than dangerous to execute a stranger's code on the company's machines.

ADR-01

Every sandbox is single-use, created for one job and destroyed after it

Accepted

Does a build environment get reused between jobs, and if so, what exactly is guaranteed to have been removed?

Context
Reusing a warm build environment is the single largest performance win available to a CI platform. A pool of containers with the dependency layers already pulled and the toolchain already warm turns a forty-second start into a two-second one, and every CI vendor has been tempted by it. The problem is that the guarantee required is not "we cleaned up the working directory". It is "nothing the previous job wrote — a file outside the workspace, a kernel object, a cached credential in an agent socket, a modified binary on the PATH, an entry in a package manager's global store — can influence or be read by this one". That guarantee is impossible to audit and its violations do not fail tests; they fail silently, in one job in ten thousand, and look like a flake.
Decision
A sandbox is created for exactly one job and destroyed after it. No sandbox serves two jobs, and the destruction happens whether the job succeeded, failed, timed out or was killed. Warmth is bought back by pre-booting sandboxes in a pool before they are claimed, never by reusing one after a job has run inside it.
How it is realised on AWS
The sandbox manager maintains a warm pool of pre-booted microVMs on each isolation host. A microVM in the pool has booted a base image and has never executed tenant code; at dispatch, one is claimed, the job's workspace and credentials are injected, and the VM is marked non-reusable. On completion the VM is terminated and its backing storage discarded, and the pool replenishes asynchronously from the host's capacity. A VM that has been claimed is never returned to the pool, even if the job never started.
Options weighed
  • ChosenSingle-use sandbox per job, warm pool of never-used sandboxes: Makes residue impossible rather than unlikely, and moves the performance problem to cold start where it is measurable
  • RejectedReused container pool with a cleanup step between jobs: Cheapest and fastest; the cleanup guarantee cannot be audited and its failures look like flakes
  • RejectedReused for trusted builds, single-use for fork builds: Assumes a trusted build cannot be hostile, which a compromised dependency disproves
  • Right elsewhereA whole dedicated instance per job: Right where jobs are long and few, or where regulatory isolation is per-customer; the cold start and cost do not survive 1.9M jobs a day
Consequences
What it buys
  • Cross-job contamination stops being a class of bug, which removes an entire category of unreproducible flake
  • The teardown path is exercised 1.9M times a day, so it is reliable by the time it matters
  • Capacity accounting becomes simple: one job occupies one sandbox for a measurable duration, which is what makes cost per job-minute meaningful
What it costs
  • Cold start becomes the platform's hardest performance problem and the warm pool a permanent cost line — up to 8% of compute minutes idle by design
  • Dependency warmth must be recovered through the cache and mirror rather than through a reused filesystem, which makes cache hit rate load-bearing
  • Cost per job-minute is materially higher than a reused-container platform's, and that gap is visible to anyone comparing vendors
Choose differently when
If every repository in the estate were internal, every contributor employed and vetted, and every dependency vendored and reviewed before entry, the residue risk would be low enough that a reused pool with a cleanup step would be the better economic choice. The decision is justified by roughly 300 fork pull requests a quarter and by a dependency graph the company does not control — remove both and the calculation changes.
Why it holds up over time
The rule is about lifecycle, not about technology. Whatever the isolation primitive becomes — a lighter hypervisor, a hardware partition, something not yet built — "one job, then destroyed" remains expressible and remains the property that makes every other isolation guarantee meaningful. The cost of the rule falls as boot times fall; the cost of abandoning it does not fall at all.
LessonWhen a guarantee cannot be audited, do not try to enforce it with a cleanup step. Change the lifecycle so the guarantee is structural, and pay for it somewhere you can measure.
Shown on views04 13 20
ADR-02

Isolation is hardware-level, and it does not vary by trust class

Accepted

Must the boundary between two concurrently running jobs survive a kernel exploit, and does the answer differ for a build of the company's own code?

Context
Containers isolate by kernel namespace, which means a kernel vulnerability reachable from inside the container is a path out of it. Syscall filtering and user-space kernels narrow that surface considerably but do not remove the shared-kernel premise. The question is not whether such exploits are common — they are rare — but what the consequence is when one lands: on a shared-kernel host, another tenant's job, its credentials and its source are in reach. The second half of the question is more interesting. It is tempting to run fork builds in microVMs and trusted builds in containers, because most jobs are trusted and containers are three times cheaper. That reasoning assumes a trusted build cannot be hostile, which has not been true since dependency compromise became the dominant supply-chain attack: a malicious postinstall script in a transitive dependency of the company's own service executes with exactly the privileges of a trusted build.
Decision
Every job runs in a hardware-virtualised microVM. Isolation strength is identical for a trusted branch build, a fork build, a scheduled run and an interactive session. What varies between trust classes is what the sandbox is given — credentials, cache write authority, publish identity — never how well it is contained.
How it is realised on AWS
Nested-virtualisation-capable EC2 instance families run a lightweight microVM hypervisor, one VM per job, with the per-job agent inside the VM and the sandbox manager outside it. Jobs that declare nested container execution get it inside their own VM without host privileges. Two instance families are qualified at all times so a capacity shortage in one is not a platform-wide throughput event.
Options weighed
  • ChosenmicroVM per job for every trust class: One substrate to harden, one cold-start number to optimise, and no capacity stranded in the wrong pool
  • RejectedHardened containers with syscall filtering for all jobs: Fastest and cheapest, and adequate until the first kernel escape, at which point it is adequate for nothing
  • RejectedmicroVMs for untrusted, containers for trusted: The common design; it prices the fork risk correctly and the dependency-compromise risk at zero
  • DeferredSeparate physical fleets per trust class: Worth revisiting if a regulator requires physical separation; it doubles the capacity-planning problem for a marginal gain over microVMs
Consequences
What it buys
  • A kernel-level compromise inside one job yields no access to another, which is the only honest answer to "what if a build roots the kernel"
  • One substrate means one hardening effort, one cold-start budget and one capacity pool — no stranded capacity in a fleet nobody is using this hour
  • Nested container execution can be offered safely, because the nesting happens inside the tenant's own VM
What it costs
  • Higher cost per job-minute and a slower start than containers, borne on every job rather than only on risky ones
  • Dependence on a narrower instance market, which is a capacity-planning risk with no software mitigation
  • Some workloads that assume host features will need explicit support rather than working by accident
Choose differently when
If the estate had no external contributors and every dependency were vendored and reviewed, the dominant residual risk would be a platform bug rather than tenant code, and hardened containers would be the right economic answer. Equally, if a future container runtime offered a genuine hardware-backed boundary at container cost, this decision becomes an implementation detail rather than a trade-off.
Why it holds up over time
The requirement is "the boundary survives a kernel exploit", which is a property, not a product. It has been satisfied by different technologies each decade and will be satisfied by others. Stating it as a property rather than as "we use microVMs" is what lets the substrate be replaced without revisiting anything else in this record set.
LessonDo not price a risk class by how often you expect its trigger. Price it by what is reachable when the trigger fires, and then refuse to grade the boundary by how much you trust the code — because the code you trust is exactly what an attacker wants to arrive inside.
Shown on views15 16 20
ADR-03

Trust is a property of the run, assigned at admission from the event

Accepted

At what moment, and from what evidence, does the platform decide whether this run's code is trusted — and can that decision ever be revised upward?

Context
Every credential and cache decision in the platform depends on one bit: is the code about to execute code the company has reviewed? Getting that bit from the wrong place is the root of most CI security incidents. If it is derived from the pipeline definition, the fork's own definition can claim to be trusted. If it is derived at credential-request time, a job that has already started can escalate. If it is recomputed per step, a step can change it. And if it is computed from the pull request's target repository rather than from the head's provenance, a pull-request trigger that runs the head's code with the base's permissions is exactly the well-documented mistake that has leaked deployment credentials from several widely-used CI systems.
Decision
The trust class is computed once, at admission, from the trigger event and the repository's membership records — never from anything inside the pipeline definition or the commit. It is written as an immutable column on the run and attached to every job, credential request and cache operation for the run's life. No path exists to raise it. Lowering it is possible only by cancelling the run.
How it is realised on AWS
The admission service reads the signed webhook payload, resolves the head repository and contributor against the tenant's membership and approval records, and writes trust_class on the run row before the run enters the queue. The job identity minted by identity federation binds tenant, repository, ref and trust class into the token, so the secret broker and cache service evaluate the class from the token rather than by looking it up again. A run's effective definition is recorded alongside the class, so an auditor can see both what ran and under what posture.
Options weighed
  • ChosenComputed once at admission from the event, immutable on the run: One place to get right, one column to audit, and no escalation path to reason about
  • RejectedEvaluated per credential request against current state: Flexible, and creates a window in which a membership change mid-run changes a running job's privileges
  • RejectedDeclared in the pipeline definition with policy validation: Lets the artefact under test participate in deciding how much it is trusted
  • Right elsewhereMaintainer-approved promotion of a fork run to trusted: Reasonable for a small open-source project where one person reviews every diff; at 610 tenants it becomes a click that nobody reads
Consequences
What it buys
  • The security posture of a run is a single auditable value, set before any capacity is committed
  • Credential and cache services need no membership lookups on the hot path, because the class travels in the token
  • The fork case becomes ordinary: it is a class, not an exception handled by a special code path
What it costs
  • A contributor who becomes trusted mid-run does not benefit until the next run, which occasionally confuses maintainers
  • The classifier is a small piece of code with disproportionate consequence, and needs test coverage out of proportion to its size
  • A legitimate need to run a fork's code with credentials has no in-platform answer, only the trusted re-run at merge
Choose differently when
If the platform served a single team with no external contributors, the class would always be the same value and the machinery would be pure overhead. And if a future source-control platform offered a cryptographically attested contributor-trust signal at event time, the classifier would become a thin adapter over it rather than a decision the platform makes itself.
Why it holds up over time
"Compute the security-relevant fact once, from the least manipulable input available, and make it immutable" is a design rule older than CI and will outlast this platform. What changes over time is which input is least manipulable; the rule does not.
LessonNever let the thing being evaluated contribute to the evaluation. Derive the security-relevant fact from the event, write it down, and give yourself no code path that can raise it later.
Shown on views05 12 15
ADR-04

The cache is read by everything and written only by trusted runs

Accepted

Who is allowed to write a cache entry that a later build will restore and execute?

Context
A build cache is not passive data. Restored into a workspace, it becomes compiler output, dependency binaries, and sometimes executables on the PATH. If an untrusted run can write an entry that a trusted run later restores, the untrusted run has achieved code execution inside the trusted build — and therefore inside whatever the trusted build is allowed to publish and deploy. This is not hypothetical: it is the most-exploited class of CI vulnerability in practice, because the cache is the one shared mutable surface a fork build legitimately touches. The counter-pressure is real too. Cache hit rate is the single largest determinant of build latency, and the contributors most in need of a fast build are exactly the ones who cannot be allowed to write.
Decision
Any run may read the cache. Only a run classified as trusted may write it. Every entry is integrity-verified on restore, and a verification failure is treated as a cache miss rather than as an error. The slower fork build is accepted as the price, and the platform tells the contributor why.
How it is realised on AWS
Cache write credentials are issued by the broker only against a job identity carrying the trusted class; an untrusted job's request is refused with a distinct, explicit reason rather than silently returning nothing. Entries are content-keyed with declared fallback keys and carry a digest verified on restore. Restore is served from a read-wide path that requires no write credential at all, so a fork build's cache read is not a privilege it could misuse.
Options weighed
  • ChosenRead wide, write privileged, integrity-verified on restore: Keeps most of the hit rate and removes the code-execution path entirely
  • RejectedShared per-repository cache writable by any branch or fork: Best hit rate, and a direct code-execution path from a stranger into the production build
  • RejectedPer-branch caches with read fallback to the default branch: Better than shared-writable, and still lets an untrusted branch's entry be read by a trusted build via fallback
  • DeferredContent-addressed cache with attested writers only: The right long-term answer and a natural extension of the attestor; deferred because it needs the remote-execution work in Phase 3
Consequences
What it buys
  • The most-exploited CI vulnerability class is removed by credential design rather than by validation
  • A poisoned or corrupt entry costs a slow build, never a compromised one, because verification failure is a miss
  • Trusted builds keep essentially the full benefit of the cache, which is where the majority of the job volume is
What it costs
  • Fork builds are measurably slower, and the contributors affected are the ones the company least wants to frustrate
  • A fork build that would have warmed the cache for its own next iteration cannot, so iteration on a fork PR is slow throughout
  • Cache hit rate must be reported per trust class, or the aggregate number hides the fork experience
Choose differently when
If the cache held only content-addressed, independently verifiable artefacts whose provenance was attested — so that restoring an entry proved who produced it and from what — then write access would stop being a privilege and any run could write safely. That is the deferred option above, and it is the direction this decision should eventually move in.
Why it holds up over time
The asymmetry between reading shared state and writing it is permanent, and it gets more valuable as caching gets more aggressive. Remote execution and distributed action caches raise the stakes rather than lowering them: the more of a build comes from cache, the more a cache write is worth to an attacker.
LessonShared mutable state between security domains is a code-execution path, whatever it is called. Split the privilege before optimising the hit rate, and be honest with the people the split slows down.
Shown on views05 14 20
ADR-05

All sandbox egress is mediated by a policy proxy, on by default

Accepted

What can a running job reach on the network, and who decides?

Context
A job needs the network: it fetches dependencies, pulls base images, clones submodules, sometimes calls a test double. Unrestricted egress makes all of that work and also makes exfiltration trivial — a build that has read something it should not can simply post it somewhere. Unrestricted egress is also how a build defeats its own recorded dependency set: a step that curls a script at build time makes the resolved-dependency record a fiction. The difficulty is that a strict allow-list breaks builds in ways that are tedious to diagnose, so any design that makes restriction opt-in results in nobody opting in.
Decision
Every sandbox's egress passes through a policy-enforcing proxy. Restriction is on by default: an untrusted run reaches only the dependency mirror and source control. A tenant may define an allow-list for its trusted runs. Every allowed and denied destination is logged and attributable to the run.
How it is realised on AWS
Isolation hosts place sandboxes on a network with no default route; the per-job agent's traffic is directed to a proxy that evaluates destination against the tenant's policy and the run's trust class. Package and image pulls are served from the platform's own mirror, which is both the performance path and the audit path. Denials return a distinct, named error so the failure reads as a policy decision rather than as a network fault.
Options weighed
  • ChosenMediated egress, restrictive by default, tenant allow-list for trusted runs: The safe posture is the one you get without doing anything, and the mirror makes it fast rather than merely safe
  • RejectedOpen egress with monitoring and alerting: Breaks nothing and prevents nothing; the alert arrives after the exfiltration
  • RejectedOpt-in allow-list per pipeline: Correct in principle; in practice almost no pipeline opts in, so the control does not exist
  • Right elsewhereNo egress at all, everything pre-provisioned into the sandbox: Right for a classified or air-gapped environment; the provisioning burden is not worth it when a mirror will do
Consequences
What it buys
  • Exfiltration from a hostile build requires defeating the proxy, not merely running code
  • The recorded dependency set becomes trustworthy, because fetching outside it is visible and usually blocked
  • Upstream registry outages degrade build latency rather than stopping all builds, because the mirror is the default path
What it costs
  • Builds that reached the internet incidentally now fail, and the first month of adoption is spent adding legitimate destinations
  • The mirror becomes a critical dependency with its own availability and storage cost
  • A determined build can still tunnel over an allowed destination; the control is a barrier, not a proof
Choose differently when
If the platform served only internal code with no untrusted runs and no supply-chain concern, the diagnostic cost of the allow-list would outweigh its benefit and monitored open egress would be defensible. As soon as one fork build runs, it is not.
Why it holds up over time
Mediating egress is independent of the network technology that implements it, and its value rises with every step the industry takes towards supply-chain attacks. The mirror that makes it palatable is also the thing that makes build reproducibility possible, so the investment pays twice.
LessonA control that must be turned on will not be. Make the restrictive posture the default and invest in making it fast, because the diagnostic pain — not the security argument — is what decides whether the control survives contact with a deadline.
Shown on views15 20 22

Trust and provenanceThe decisions that make the link between commit, build and deployed artefact non-forgeable.

ADR-06

The attestor signs provenance; the build never signs anything

Accepted

Who makes the statement "this artefact was built from this commit by this pipeline", and why should anyone believe it?

Context
A provenance attestation is only worth the difficulty of forging it. If the signing key is available inside the build — as an environment variable, a mounted file, or a credential the job can exchange — then the statement is made by the thing being described, and a compromised build can produce an artefact together with a perfectly valid attestation saying it came from somewhere else. That is not a marginal weakness; it makes the entire evidence chain decorative, because the case the chain exists to detect is precisely a compromised build. The counter-argument is practical: the build knows most of the facts, so having it assemble and sign the statement is much simpler than reconstructing them outside.
Decision
The signing key is held by a control-plane attestor and is never present inside a sandbox. The job reports its outcome and uploads its artefact; the attestor composes the attestation from the inputs the platform itself observed — the resolved effective definition, the source commit, the recorded dependency set, the builder identity and the timings — and signs it. The build's own claims are inputs to be recorded, never assertions to be signed.
How it is realised on AWS
The signing key lives in a hardware-backed key store reachable only by the attestor's role; no path exists from an isolation host to it. The attestor is invoked by the sandbox manager after the job reports, binds the request to the job's lease and identity, composes the statement from control-plane records, signs, and appends to the transparency log. The artefact is addressed by the digest of its bytes, so the statement is about specific bits rather than about a name.
Options weighed
  • ChosenControl-plane attestor signs from observed inputs: A compromised build can produce a bad artefact but not a credible claim about one
  • RejectedThe build signs with a short-lived key issued to the job: Much simpler, and the statement is then made by the code under suspicion
  • RejectedThe build composes the statement, the control plane signs it unexamined: Looks like the chosen option and is the rejected one: the content is still the build's
  • DeferredA separate rebuilder independently reproduces and signs: The strongest available evidence and roughly doubles build cost; worth it for a small set of release artefacts, not for 1.9M jobs a day
Consequences
What it buys
  • An attestation is worth verifying, which is the entire point of producing one
  • Verification can be performed outside the platform against the transparency log, so the platform is not the only witness to its own claims
  • The gate service has something objective to check before a deployment, rather than a self-report
What it costs
  • The attestor must reconstruct facts the build knew directly, which means the control plane has to record more than it otherwise would
  • The attestor becomes the highest-consequence code in the platform: a confused-deputy bug there yields valid, false statements
  • Signing is on the critical path of every job completion, adding latency and a dependency to the seal step
Choose differently when
If artefacts were never deployed anywhere consequential — a research estate, a throwaway environment — the cost of a control-plane attestor would exceed its value and a build-signed statement would be adequate record-keeping. The decision is justified by the artefact reaching production under a gate that trusts the statement.
Why it holds up over time
"The observer signs, not the observed" predates software supply chains and will outlive every signing format. Sigstore, in-toto, whatever replaces them: each is a way of expressing this separation, and the separation is the part worth committing to.
LessonEvidence produced by the subject of the evidence is not evidence. If a claim must survive the compromise of the thing it describes, the claim has to be made somewhere else.
Shown on views02 13 20
ADR-07

Credentials are minted per job and revoked when the job ends

Accepted

What credential is inside a running job, how long is it useful, and what happens to it when the job stops?

Context
A long-lived credential inside a build environment is the most valuable thing an attacker can obtain from a CI platform, because it outlives the build, works from anywhere, and usually has more authority than the build needed. Static secrets in a pipeline's configuration are the classic form; a cloud role attached to the runner host is the modern one, and it is worse, because every job on that host inherits it whether it needed it or not. The alternative — minting an identity per job and exchanging it for narrowly scoped, short-lived credentials — puts identity federation on the hot path of every one of 1.9M daily jobs, which is a real availability and latency cost.
Decision
Every job is issued a workload identity binding tenant, repository, ref and trust class, valid for the job's lifetime. All secret and cloud access is scoped to that identity. Nothing long-lived is present in a sandbox. Revocation is pushed at job end rather than left to token expiry, and every secret read is recorded in the audit trail before the value is returned.
How it is realised on AWS
The scheduler asks identity federation to mint the job identity at dispatch; the token is delivered to the per-job agent and never to the job's own environment except as an explicitly requested, scoped exchange. The secret broker evaluates the token's trust class against policy, writes the audit record, then returns the value with redaction registered on the log path. Cloud access is obtained by exchanging the identity with the cloud's token service for a short-lived role. At job end the scheduler revokes the identity rather than waiting for the TTL.
Options weighed
  • ChosenPer-job identity, scoped exchange, revoked at job end: A leaked credential's useful life is the job's life, and its authority is the job's authority
  • RejectedHost-attached instance role shared by every job on the host: Simplest and requires no federation; every job inherits authority it did not ask for
  • RejectedStatic secrets in pipeline configuration: Still the industry's most common design and the source of most credential leaks
  • DeferredPer-step identity with per-step scope: Strictly better and multiplies federation traffic by the step count; revisit if step-level blast radius becomes the binding concern
Consequences
What it buys
  • A credential exfiltrated from a build is worthless minutes later and was never broadly scoped
  • Every secret read has an audit record naming the job, the repository, the ref and the trust class
  • Denying an untrusted run is a policy evaluation on a token rather than a special case in the build path
What it costs
  • Identity federation is on the hot path of every job and must be more available than the execution plane it serves
  • Token minting and exchange add latency at job start, inside an already-tight cold-start budget
  • Pipelines that expect a static secret must be migrated, and the migration is visible to every team
Choose differently when
If the platform's jobs needed no external authority at all — pure compile-and-test with no publish and no cloud access — the federation machinery would be unnecessary and a host role with no permissions would do. The moment a job can publish or deploy, it does not.
Why it holds up over time
Short-lived, narrowly scoped, workload-bound credentials are the direction every identity system has moved in for fifteen years and the direction the cloud providers keep extending. Building on federation rather than on stored secrets means the platform inherits those improvements instead of having to migrate to them.
LessonMeasure a credential by how long it is useful after it leaks and by how much it can do that its holder did not need. Both numbers should be as close to zero as the platform's latency budget allows.
Shown on views13 20 21
ADR-08

Attestations are recorded in an append-only log that can be verified without the platform

Accepted

If the platform itself were compromised or simply wrong, how would anybody find out that an attestation had been altered or invented?

Context
An attestation stored in a mutable database beside the run metadata is trustworthy exactly as far as the platform is trustworthy. That is a circular guarantee, and it fails in the two cases evidence is most needed: a compromise of the platform's own control plane, and an honest bug that mis-attributes one job's inputs to another's artefact. Both produce records that are internally consistent and factually wrong, and neither is detectable from inside. The cost of doing better is real: an append-only, tamper-evident log is more operationally awkward than a table, cannot be corrected in place, and forces the platform to keep records it might prefer to tidy.
Decision
Every attestation is appended to a tamper-evident, append-only log, retained seven years, verifiable by a third party without the platform's cooperation. Nothing in the log is ever edited or deleted. An artefact whose provenance fails verification is quarantined rather than removed, because a deleted artefact cannot be investigated.
How it is realised on AWS
Attestations and SBOMs are written to object storage under an immutability policy, with the log's inclusion proofs published so an external verifier can check an entry without asking the platform to vouch for it. The transparency log sits in the custody zone, reachable from the attestor and from nothing in the execution plane. Cross-region replication carries evidence at RPO 0, because losing proof is worse than losing work.
Options weighed
  • ChosenAppend-only tamper-evident log, externally verifiable, 7 years: Makes an inconsistency discoverable without the platform's cooperation, which is the only case that matters
  • RejectedAttestations as rows beside run metadata: Simple, queryable, and worthless in the two failure cases the evidence exists for
  • Right elsewherePublic transparency log shared with the wider ecosystem: Right for open-source artefacts the world consumes; for internal artefacts it leaks the shape of the estate
  • RejectedPeriodic signed snapshots of the attestation table: Cheaper, and detects tampering only at snapshot granularity, which an attacker chooses to work inside
Consequences
What it buys
  • An auditor can answer "which commit produced this image" without the platform being a trusted intermediary
  • A control-plane bug that mis-signs becomes discoverable after the fact rather than invisible
  • Quarantine rather than deletion keeps the investigation possible, and the artefact still cannot deploy
What it costs
  • Nothing can be corrected in place; an error is fixed by appending a correction, which is operationally unfamiliar
  • Seven years of immutable evidence is a storage and legal-hold commitment that grows monotonically
  • The immutability policy also protects records the company might later wish to remove, which is the point and is sometimes inconvenient
Choose differently when
If artefacts were short-lived and never subject to audit, retention of any kind would be waste and a table would suffice. The decision is driven by two audit reviews a year and by production images whose origin must be provable years after the engineer who built them has left.
Why it holds up over time
Tamper-evident append-only logs have outlived several generations of signing technology, and the reason is structural: the guarantee comes from the data structure rather than from the operator. Whatever replaces today's log formats will still be a log, and its entries will still need to be verifiable by someone who does not trust the writer.
LessonEvidence held only by the party it exonerates is not evidence. Make the record append-only and externally checkable, and accept that you can no longer tidy it.
Shown on views11 20 21

Fairness and capacityThe decisions that stop one tenant's burst becoming everyone else's queue.

ADR-09

Fairness is enforced in the scheduler, by entitlement and weighted fair queuing

Accepted

When 40,000 slots are full and one tenant has just pushed a 900-job monorepo pipeline, whose job runs next?

Context
A global FIFO queue with per-tenant concurrency caps is the usual first design and it is comprehensible, cheap and wrong in one specific way: a cap bounds how much capacity a tenant holds but not how much of the queue it occupies. A monorepo pipeline that admits 900 jobs at once puts 900 entries ahead of the next tenant's five, so the small tenant's wait is set by the large tenant's burst regardless of the cap. Because CI latency is felt personally — an engineer is sitting there — the small tenant experiences this as the platform being broken, and the platform team experiences it as an email asking large teams to be considerate. Weighted fair queuing fixes it, at the cost of needing cross-tenant state in the dispatch decision.
Decision
Each tenant has a concurrency entitlement and a priority weight. Dispatch selects across tenants by weighted fair queuing rather than by arrival order, so queue wait inside an entitlement is bounded independently of any other tenant's burst. Idle entitlement is lent to tenants that are over theirs, and reclaimed when the owner returns. Queue wait inside entitlement is an SLO; a tenant over its entitlement has no wait guarantee.
How it is realised on AWS
The scheduler is sharded, with per-shard fair-queue state and a leader per shard, and entitlements held in the run metadata store. Priority tiers — interactive, scheduled, background — are evaluated before fairness, so an interactive job outranks a background job of any tenant, and background work is preemptible. Queue position and estimated wait are computed from the same state and exposed to the waiting engineer.
Options weighed
  • ChosenPer-tenant entitlement with weighted fair queuing and lending: Bounds the small tenant's wait and still lets a big tenant use idle capacity
  • RejectedGlobal FIFO with per-tenant concurrency caps: Simple and cheap; the small tenant's wait is set by the large tenant's burst
  • RejectedHard per-tenant partitions, no lending: Perfectly fair and strands capacity inside unused entitlements, which at 610 tenants is most of the fleet most of the time
  • DeferredA market with priced priority: Aligns incentives elegantly and requires a chargeback culture the organisation does not yet have; revisit alongside showback
Consequences
What it buys
  • A two-person team's five-job pipeline has a bounded wait no matter what the monorepo is doing
  • Idle entitlement is usable, so fairness does not cost utilisation
  • Queue position and estimated wait become computable, which turns an unexplained wait into an explained one
What it costs
  • The scheduler needs cross-tenant state in the dispatch decision, which constrains how far it can be sharded
  • Entitlements become a thing platform engineers must curate, and a badly-set entitlement is now the cause of a wait
  • Fairness violations need their own monitoring, because the failure is silent from any single tenant's point of view
Choose differently when
If tenant sizes were roughly uniform and bursts rare, FIFO with caps would deliver nearly the same outcome for a fraction of the complexity. The decision is justified by a distribution in which a handful of monorepos generate a large share of job volume — and that distribution is getting more extreme, not less.
Why it holds up over time
Weighted fair queuing is fifty-year-old theory from packet scheduling and has survived every change of substrate since. The pressure it addresses — the ratio between the largest and smallest workload sharing a resource — increases with monorepo size, generated code volume and AI-assisted development, so the decision gets stronger with time.
LessonA concurrency cap limits what a tenant holds, not what a tenant blocks. If wait time is the experience, fairness has to be a property of the queue, not of the quota.
Shown on views02 03 18
ADR-10

The warm pool is tenant-agnostic, and a claimed sandbox is never returned to it

Accepted

Can a sandbox be pre-warmed with a tenant's image and dependency layers, and if so, may it later be handed to a different tenant?

Context
Single-use sandboxes (ADR-01) make cold start the platform's hardest performance problem, and the obvious mitigation is to pre-warm with tenant content: pull the base image, hydrate the dependency layers, then hand the ready sandbox to that tenant's next job. The saving is large — most of the queued-to-first-step budget is image pull and dependency hydration. The problem is that a pre-warmed sandbox is now tenant-specific, so it can only serve that tenant; at 610 tenants with uneven arrival rates, most pre-warmed sandboxes sit idle waiting for a job that does not come, and the ones that are needed are in the wrong pool. Handing a tenant-warmed sandbox to another tenant is the alternative, and it reintroduces exactly the residue problem ADR-01 exists to remove.
Decision
The warm pool holds tenant-agnostic sandboxes that have booted a base image and have never seen tenant content. One is claimed at dispatch, the workspace and credentials are injected, and it is never returned to the pool. Tenant-specific warmth is recovered through the cache and the dependency mirror, not through a pre-warmed sandbox.
How it is realised on AWS
Each isolation host maintains a small pool of pre-booted microVMs from a common base image, sized from the observed arrival rate for that host's resource class. Dispatch claims one and marks it non-reusable. Tenant warmth comes from the cache restore path and from the mirror serving image and package layers from the same region, which is where the remaining cold-start budget is spent and optimised.
Options weighed
  • ChosenTenant-agnostic pre-booted pool, claimed at dispatch, never returned: One pool serves every tenant, so the idle-waste ceiling is a single number to manage
  • RejectedPer-tenant pre-warmed pools with tenant image layers: Best cold start; at 610 tenants most warm capacity is idle in the wrong pool
  • RejectedTenant-warmed pools that may be reassigned between tenants: Has the cold-start benefit and reintroduces the residue problem ADR-01 removes
  • DeferredPer-tenant pools for the largest tenants only, agnostic for the rest: Defensible hybrid worth revisiting once per-tenant arrival rates are measured rather than assumed
Consequences
What it buys
  • Idle warm capacity is one fungible pool, so the 8% waste ceiling is a number that can actually be held
  • No sandbox ever holds one tenant's content while waiting for another tenant's job, so ADR-01 survives intact
  • The cache and mirror become the levers for cold start, and both benefit every tenant rather than the one they were warmed for
What it costs
  • Cold start is longer than a tenant-warmed pool would give, and that difference is the most visible performance number the platform publishes
  • Cache hit rate and mirror locality carry more weight than they otherwise would, which makes them availability-critical
  • A very large tenant with unusual images gets no special treatment, and will ask for it
Choose differently when
If the estate were a handful of large tenants with steady arrival rates and huge images, per-tenant pools would be both affordable and materially faster, and the hybrid above becomes the right answer. The decision is driven by 610 tenants with uneven arrival.
Why it holds up over time
The trade — fungible capacity against specific warmth — is a permanent one in any pooled system, and the right side of it is decided by the number of tenants and the variance of their arrival. Both of those are measurable, so this decision is designed to be revisited with data rather than with argument.
LessonPre-warming with tenant-specific content converts a capacity problem into a placement problem. Before paying for the warmth, check whether the pool will be in the right place when the job arrives.
Shown on views13 16 18
ADR-11

Interruptible capacity is confined to the background tier

Accepted

Which jobs may run on capacity the cloud can take back, and what happens when it does?

Context
Interruptible capacity is substantially cheaper and CI is a near-ideal workload for it: jobs are short, stateless and retryable. The temptation is to put everything on it and re-dispatch on reclamation. That works arithmetically and fails experientially, because a reclamation adds the full cold start plus the elapsed work to a job someone is sitting and waiting for, and the distribution of that added latency is exactly the long tail the queue-wait SLO exists to bound. Meanwhile the nightly and backfill work, which nobody is waiting for, is the natural home for the risk.
Decision
Interruptible capacity serves the background tier only: scheduled runs, nightly work, backfills. Interactive and scheduled-but-blocking work runs on on-demand capacity. A reclaimed job is re-dispatched to on-demand capacity, and the drain notice is honoured where the provider gives one. The target is that at least 55% of eligible background job-minutes run on interruptible capacity.
How it is realised on AWS
Resource classes carry an eligibility flag; the scheduler places eligible background jobs on interruptible hosts and everything else on on-demand. The per-job agent handles the drain notice by reporting a distinct platform-failure cause and terminating cleanly, so the re-dispatch is attributable and does not pollute the infrastructure-failure ratio's interpretation. Background work is also preemptible by interactive work on shared hosts.
Options weighed
  • ChosenInterruptible for background tier only, re-dispatch to on-demand: Takes most of the saving without putting the waiting engineer's tail latency at risk
  • RejectedInterruptible for everything with transparent re-dispatch: Best unit cost and worst p99 for exactly the jobs someone is watching
  • RejectedOn-demand only: Simplest capacity story and leaves a large, easily-harvested saving on the table
  • Right elsewhereInterruptible with checkpoint and resume mid-job: Right for long ML training jobs; a 3-minute CI job is cheaper to restart than to checkpoint
Consequences
What it buys
  • A meaningful share of compute cost is harvested from work nobody is waiting for
  • The interactive queue-wait SLO is not exposed to reclamation tail latency
  • Reclamation becomes an ordinary, attributable event rather than an incident
What it costs
  • Nightly throughput is variable, so a capacity-scarce night lengthens the nightly window
  • Two capacity pools to plan and monitor rather than one
  • The eligibility flag is a decision someone has to make per resource class, and a wrong one is felt as latency
Choose differently when
If the interactive volume were small and latency-insensitive — a batch-oriented estate, or one where CI results are read the next morning — putting everything on interruptible capacity would be correct. Equally, if reclamation rates in the chosen families fell far enough, the distinction would stop earning its complexity.
Why it holds up over time
The rule generalises past any particular cloud's interruptible product: place cheap-but-unreliable capacity where the latency distribution does not matter. That mapping between workload patience and capacity reliability is permanent even as the products change names.
LessonCheap capacity is not free capacity; it is paid for in tail latency. Spend that tail on work with no one waiting at the end of it.
Shown on views16 18 22

Run state and recoveryThe decisions that make a partition resolve to re-dispatch rather than to a false green.

ADR-12

The relational store is authoritative for scheduling; the event log is authoritative for history

Accepted

When the control plane restarts mid-run, what does it read to find out what was happening — and if two sources disagree, which one wins?

Context
Run state has two consumers with incompatible needs. Scheduling needs a strongly consistent, transactional view it can make decisions against right now: is this job leased, is this tenant at its entitlement, may this run proceed. History, recovery and audit need an ordered, immutable account of everything that happened, including the transitions a current-state row has already overwritten. Serving both from one store means either a relational table that loses history, or an event log that cannot answer "how many slots is this tenant holding" without a fold over everything. Serving them from two stores means they can disagree, and a design that does not say which one wins will discover the answer during an incident.
Decision
Run and job current state live in a transactional relational store, which is authoritative for every scheduling decision. Every transition is also appended to an ordered event log, which is authoritative for history, replay and audit. On disagreement, the relational store wins for scheduling and the event log wins for what happened. The log is retained independently, so losing the metadata store does not lose the account.
How it is realised on AWS
Aurora PostgreSQL holds runs, jobs, leases and entitlements with RPO 0 and RTO 15 min; transitions are appended to a Kinesis stream with RPO 0 and RTO 30 min, retained 400 days. The scheduler reads and writes the relational store transactionally and emits to the log as part of the same commit path, so a transition cannot be applied without being recorded. Recovery rebuilds the in-memory scheduler view from the relational store, then reconciles against the log to find transitions that were recorded but not applied.
Options weighed
  • ChosenRelational authoritative for scheduling, event log for history, stated precedence: Each consumer reads the store that fits it, and the tie-break is written down before it is needed
  • RejectedEvent log as the single source of truth, state folded on read: Elegant and immutable; folding for every dispatch decision at 3,200 jobs a minute is the wrong hot path
  • RejectedRelational only, history in an audit table: Simpler; the audit table is mutable and shares the store's failure domain, so it fails when it is needed
  • RejectedThe runner's own report as authoritative for its job: Removes a store and makes a partition indistinguishable from a success — see ADR-13
Consequences
What it buys
  • Dispatch decisions are transactional and fast, and history is complete and immutable
  • Recovery has a defined procedure rather than an argument, because precedence was decided in advance
  • Losing the metadata store loses current state, not the account of what happened
What it costs
  • Two stores to operate, replicate and reason about, with a reconciliation step in the recovery path
  • Every transition costs a write to both, which is a throughput floor on the whole platform
  • A bug that writes one and not the other produces a divergence that only the reconciliation finds
Choose differently when
At a fraction of this volume, folding the log on read would be entirely affordable and the second store unnecessary. Conversely, if scheduling state grew large enough that a single relational store became the bottleneck, the answer is sharding it rather than collapsing back to one store.
Why it holds up over time
The separation of current state from the account of how it got there is a distinction, not a technology, and it survives every substitution of database or stream. Naming the precedence rule is the part that pays off, because that is what an on-call engineer needs at 03:00 and cannot derive.
LessonTwo stores are fine. Two stores without a written precedence rule are one incident away from being a guess.
Shown on views11 12 19
ADR-13

Runner loss is detected by lease expiry, and a job is never successful on an absent runner

Accepted

A job has stopped reporting. Is it dead, slow, or finished successfully with a lost response — and what does the platform do about it?

Context
This is the network-partition question in its most concrete form, and CI platforms get it wrong in a characteristic way: they treat the absence of a failure report as evidence of success, or they treat the absence of a heartbeat as evidence of death and re-dispatch a job that is in fact still running. Both are bad, but they are not equally bad. Re-dispatching a live job costs capacity and, for a job with side effects, may duplicate them. Marking an absent job successful puts an unverified green check on a pull request, which is the failure this platform exists to prevent. The asymmetry decides the design.
Decision
Every dispatched job holds a lease that it must renew. Loss is detected by lease expiry, never by silence on the result channel. An expired lease means the job is re-dispatched with the attempt count incremented and the cause recorded. A job is never reported successful on the basis of an absent runner. Automatic retry happens only for platform failure or a declared-retryable infrastructure error, never for user failure, and every retry records its cause.
How it is realised on AWS
The scheduler issues a lease with each job assignment and the per-job agent renews it; expiry re-queues the job. Partial logs already streamed are preserved and attached to the failed attempt, so the engineer sees what happened before the loss. Because the artefact is sealed by the attestor after the report, a job whose report is lost has no attestation and therefore cannot be treated as complete even if its bits were uploaded.
Options weighed
  • ChosenLease expiry detects loss; never successful on absence; retry only platform failure: Resolves a partition to re-dispatch, which is the safe side of the asymmetry
  • RejectedHeartbeat timeout with optimistic completion on late report: Reduces wasted work and occasionally produces a green check nobody earned
  • RejectedTrust the runner's report, no lease: Simplest; a silent runner becomes indistinguishable from a passing build
  • RejectedRetry every failure a fixed number of times: Hides real failures, wastes capacity on compile errors, and makes the infrastructure-failure ratio unmeasurable
Consequences
What it buys
  • A partition costs duplicated work rather than a false green, which is the correct side to fail on
  • Retry cause is recorded, which is what makes the infrastructure-failure ratio meaningful rather than decorative
  • Partial logs survive a lost runner, so the engineer has something to read
What it costs
  • A live-but-partitioned job is re-dispatched and its work duplicated, occasionally twice
  • Jobs with external side effects need idempotency of their own; the platform cannot provide it for them
  • Lease renewal is traffic on the hot path of every job, and its own availability matters
Choose differently when
If jobs were long, expensive and side-effecting — a multi-hour deploy rather than a three-minute test — the cost of duplicating one would exceed the cost of waiting longer, and a design with a longer lease plus explicit operator adjudication would be better. At a median job of 3m 10s, re-dispatch is the cheap option.
Why it holds up over time
"Absence of evidence is not evidence of success" is a distributed-systems invariant, and leases have been the mechanism for expressing it for decades. What changes is the timeout tuning; the asymmetry that decides the design does not.
LessonWhen you cannot distinguish two states, choose which error you would rather make and design for it explicitly. Here, wasting a build is always better than blessing one.
Shown on views13 19 22
ADR-14

Six job outcomes, and a platform failure is never reported as the user's

Accepted

What does a red check actually mean, and can the engineer tell from it whether the problem is theirs?

Context
Most CI systems report two outcomes: pass and fail. Everything that is not a passing test — an image pull timeout, a reclaimed instance, a registry outage, a cancelled run, a policy denial — arrives as fail, and the engineer's only recourse is to read the log and guess. The consequence is cultural rather than technical. An engineer who cannot distinguish a flake or an infrastructure blip from their own regression learns that the cheapest response to red is to press retry, and after that they retry genuine failures too. At that point the platform has trained its users out of reading its output, and no amount of dashboard investment recovers it.
Decision
A job reports exactly one of six outcomes: success, user failure, platform failure, cancelled, timed out, or blocked by policy. A platform failure is never attributed to the user's code. The infrastructure-failure ratio — platform failures as a share of all jobs — is a headline SLO at ≤ 0.5% weekly, alerting at 1.0% hourly. Flake rate, defined as differing outcomes on an identical commit and definition digest, is measured and reported per pipeline but deliberately not bounded: detection is the platform's obligation, fixing the test is the owning team's.
How it is realised on AWS
The per-job agent and the sandbox manager classify at the point of failure, where the cause is known, rather than post-hoc from a log. Retry policy keys off the class. The run console surfaces the class, the cause and a first-failure summary — the failing job, step and the log region around the first error. Flake detection compares outcomes across attempts on an identical commit and definition digest.
Options weighed
  • ChosenSix outcomes, retry keyed off class, infra-failure ratio as a headline SLO: Makes trust in a red build measurable, and makes retry a decision rather than a reflex
  • RejectedPass/fail with the cause in the log: Universal and cheap; it is how platforms teach engineers to stop reading
  • RejectedPass/fail/error, three outcomes: Better, and still conflates cancellation, timeout and policy denial with infrastructure error
  • DeferredAutomatic flake quarantine after N differing outcomes: Attractive and hazardous: quarantining a test is how a real regression ships. Report it, do not disable it
Consequences
What it buys
  • An engineer can tell from the outcome alone whether the problem is theirs, which is what keeps a red build worth reading
  • The infrastructure-failure ratio becomes measurable, and it is the number that decides whether the platform is trusted
  • Retry stops being a cultural reflex and becomes a classified, recorded action
What it costs
  • Every failure path in agent and manager must classify correctly, which is real work spread across the whole codebase
  • Six outcomes are more than a badge can express, so the surfaces that show them need design attention
  • A misclassification is worse than no classification, because it is confidently wrong
Choose differently when
If the platform served a single team that reads every log anyway, the taxonomy would be documentation rather than mechanism. It earns its cost at 4,200 engineers who will never read the platform's documentation and will absolutely learn a reflex.
Why it holds up over time
The distinction between "your change is wrong" and "our platform failed you" is about accountability, not technology, and it will matter identically on any future substrate. Designing the outcome taxonomy so the ratio is measurable is a permanent asset that survives every migration.
LessonThe output of a build system is not a boolean; it is an attribution. Get the attribution wrong and users stop reading the output, which costs more than any latency regression.
Shown on views04 18 22

Promotion and gatingThe decisions that make "did we ship what we tested" answerable, and the gates unweakenable by the author.

ADR-15

Build once, promote the digest, bind configuration at deployment

Accepted

Does the artefact that reaches production come from the run that passed staging, or from a fresh build of the same commit?

Context
Rebuilding per environment is intuitive and common: the same commit, built again with the environment's settings, gives an artefact tailored to its target. It also breaks the only chain of evidence linking what was tested to what is running, because the production artefact has never been tested — an artefact built from the same source is not the same artefact unless the build is bit-for-bit reproducible, and almost no real build is. Timestamps, dependency resolution drift, base-image tag movement and toolchain patch levels all vary between two builds minutes apart. The counter-pressure is that build-time configuration is genuinely convenient, and teams will have baked environment assumptions into their images for years.
Decision
One artefact digest is produced once and promoted, unchanged, through every environment. Environment-specific configuration is bound at deployment time, not at build time. Promotion moves a pointer; the bits never change. Rollback repoints to a retained previously-deployed digest and is an ordinary operation with a latency target rather than a recovery procedure.
How it is realised on AWS
Artefacts are content-addressed, and an environment holds a pointer to one digest plus its configuration and gates. The promotion path is ordered per pipeline, and skipping a stage requires a recorded break-glass override with actor, reason and expiry. A deployment lock serialises promotions to one environment. Retention keeps prior production digests so rollback has somewhere to point.
Options weighed
  • ChosenBuild once, promote by digest, configuration at deploy time: "Did we ship what we tested" gets a cryptographic answer, and rollback becomes a pointer move
  • RejectedRebuild per environment from the same commit: Convenient for build-time configuration, and the production artefact has never been tested
  • RejectedBuild once, then re-tag and re-sign per environment: Keeps the bits and adds a mutable signing step that weakens the provenance for no gain
  • DeferredReproducible builds so a rebuild is provably equivalent: Would make rebuilding safe and is a multi-year effort across every language in the estate; a good ambition, not a plan
Consequences
What it buys
  • The production artefact is byte-identical to the tested one, and the attestation covers exactly those bits
  • Rollback is a five-minute pointer move rather than a twenty-minute rebuild under pressure
  • Audit becomes a query: one digest, one attestation, one chain of gate evaluations
What it costs
  • Build-time configuration must be migrated to deployment-time binding, which is visible work for every team
  • Prior digests must be retained per environment so rollback has a target, which is storage the platform pays for
  • One artefact must be able to run in every environment, which constrains how images may be built
Choose differently when
If builds were genuinely bit-for-bit reproducible across the whole estate, rebuilding per environment would be provably equivalent and the constraint on build-time configuration could be relaxed. That is the deferred option, and it is the only condition that changes this decision.
Why it holds up over time
Content addressing plus a mutable pointer is the pattern every artefact ecosystem has converged on, from container registries to package managers to source control itself. It outlives any particular registry because the identity comes from the bytes rather than from the store.
LessonIf two things must be the same, do not build them twice and compare. Build once and move the result — and then the sameness needs no argument.
Shown on views06 14 17
ADR-16

Gates are evaluated outside the pipeline they govern, and fail closed in production

Accepted

Can the author of a pipeline weaken the checks that decide whether their own change reaches production?

Context
Gate jobs inside the pipeline are the natural design: the checks live beside the code they check, they are versioned with it, and a team can evolve them. They are also self-certification. The pipeline definition is authored by the same people whose change is being gated, so a gate expressed as a pipeline job can be reordered, made conditional, marked continue-on-error, or removed — usually with an entirely sincere reason, under deadline pressure, in a diff nobody reviews closely because it is 'just CI config'. The second half of the question is what happens when the external decision service is unavailable: failing open keeps delivery moving and silently removes every control, and failing closed stops delivery during an outage of something that is not the application.
Decision
Gates are evaluated by a decision service outside the pipeline, consulted by the deployer. A pipeline author cannot weaken the gates applying to their deployment. The service returns allow, deny, or allow-with-override-required, always with a machine-readable reason naming the failing condition. Production fails closed when the service is unavailable; a tenant may configure fail-open for non-production only. Overrides are recorded with actor, reason and expiry, and override usage is a watched signal.
How it is realised on AWS
Policy bundles are versioned per repository and held in the control plane, cached per tenant for a p99 decision under 500 ms. The deployer consults the service before moving an environment pointer; the verdict and its reason are written to the immutable audit trail before the pointer moves. Gate types include required job outcomes, code-owner and human approval, provenance and signature verification, vulnerability-severity thresholds, change-freeze windows, target SLO health and a required soak. Separation of duties is enforced where tenant policy demands it.
Options weighed
  • ChosenExternal decision service, fail closed in production, recorded overrides: The only design in which a gate is not under the control of the change it gates
  • RejectedGate jobs inside the pipeline: Simplest and most flexible, and is self-certification wearing an automation costume
  • Right elsewhereGates in the target platform's admission layer: Excellent for runtime invariants and blind to build-time evidence; complementary rather than alternative
  • RejectedExternal service with fail-open everywhere: Keeps delivery moving during an outage by removing every control at exactly the moment nobody is watching
Consequences
What it buys
  • A gate means the same thing on every pipeline, and changing it is a policy change with an owner
  • Every verdict, approval and override has an audit record attributable to an identity
  • Provenance verification becomes enforceable rather than advisory, because the enforcer is not the build
What it costs
  • Every promotion depends on a service that must be more available than the deployments it governs
  • Teams lose the ability to express a local, sensible gate as a pipeline job and must go through policy
  • Break-glass override is necessary and is the control most likely to be abused
Choose differently when
If every team were both the author and the accountable owner of its production environment, with no separation-of-duties requirement and no shared blast radius, in-pipeline gates would be adequate and much simpler. The decision is driven by nine business units sharing a platform and by audit obligations that require the gate to be outside the gated.
Why it holds up over time
Separating the decision from the thing being decided is a governance invariant, and the specific gate types will churn constantly — vulnerability thresholds, freeze windows, SLO health, whatever comes next — without disturbing the boundary. That is exactly why the boundary rather than the gate list is what this record commits to.
LessonA control the controlled party can edit is not a control. Put the decision somewhere the change cannot reach, and make the override expensive to use rather than impossible.
Shown on views06 17 18

Cost, retention and developer trustThe decisions that keep the safe path the cheap path, and a red build worth reading.

ADR-17

Retention is set by artefact provenance class and pipeline class, not globally

Accepted

How long is a build artefact, a log and an attestation kept — and is the answer the same for a pull-request build as for a release?

Context
A global retention policy is easy to state and always wrong in one of two directions. Set it long enough for audit and the platform pays to keep 480 TB of pull-request artefacts and 6 TB a day of logs from builds nobody will ever look at again. Set it short enough to be affordable and the evidence an auditor needs is gone. The volumes make this the platform's fastest-growing cost line and the one most easily wasted, because nobody notices over-retention until the bill arrives, and by then the data cannot be un-kept cheaply.
Decision
Retention is a property of what produced the data. Release artefacts are kept 400 days, main-branch artefacts 90 days, pull-request artefacts 14 days. Logs are searchable 14 days, retrievable 90 days, and archived 400 days for release pipelines only. Caches are evicted after 7 days idle. Provenance attestations, the transparency log and the audit trail are kept 7 years and are immutable. Run metadata is kept 400 days. Eviction is auditable, and nothing referenced by a live environment pointer or a retained attestation is ever evicted.
How it is realised on AWS
Retention class is written on the artefact at seal time from the run's trigger and ref, so eviction never has to re-derive intent later. Object lifecycle policies implement the tiers; the immutable classes sit under an object-lock policy in the custody zone. Cost per job-minute and artefact TB by class are reported per tenant so the expensive habits are visible to the teams that have them.
Options weighed
  • ChosenRetention by provenance and pipeline class, immutable evidence separate: Spends storage where it is asked for and reclaims it where it never will be
  • RejectedOne global retention period for everything: Either unaffordable or non-compliant, and usually both at different layers
  • RejectedPer-tenant retention chosen by each team: Every team chooses the maximum, because nobody pays for storage they do not see
  • DeferredUsage-driven retention — keep what has been accessed recently: Efficient for artefacts and dangerous for evidence, which is by definition accessed rarely and needed absolutely
Consequences
What it buys
  • The largest volumes are the shortest-lived, so the cost line tracks activity rather than accumulating forever
  • Audit obligations are met by a small, immutable, clearly-bounded set of data
  • Retention class is decided at seal time from facts, so eviction never needs to guess at intent
What it costs
  • More classes to reason about, and a wrongly-classed artefact is either wasteful or gone too soon
  • Pull-request artefacts vanish in 14 days, which occasionally frustrates a long-running investigation
  • The 7-year immutable set cannot be trimmed, so a classification error there is permanent
Choose differently when
If artefact and log volumes were small enough that storage was not a material cost, a single long global retention would be simpler and harmless. At 480 TB growing 40% a year and 6 TB of logs a day, it is the fastest way to waste money on data nobody wants.
Why it holds up over time
Tying retention to the provenance of the data rather than to a global default is a policy shape, not a storage feature, and it transfers to any storage technology. As volumes grow, the argument for it strengthens; it never weakens.
LessonRetention is not one number. Decide it from what produced the data, write the class down at the moment you still know why, and make the cost visible to whoever generates it.
Shown on views10 11 18

Every package used, in one table

The terms below are used precisely in this package. Several are used loosely in the wider CI/CD literature, and the difference matters when reading the decision records.

PackageWhat it isWhat it does hereConsidered instead
Run One execution of a pipeline definition for one trigger event, holding a DAG of jobs. The unit of admission, trust classification and audit. A 'build', which in common usage means either a run or a single job and so is avoided here.
Job One node of a run's DAG, executed in exactly one sandbox. The unit of scheduling, isolation, credential scope and outcome reporting. A 'step', which is a command inside a job and shares neither its sandbox boundary nor its outcome.
Trust class Trusted or untrusted, computed at admission from the trigger event and membership records. The single bit that decides credentials, cache write authority and publish identity. A permission declared in the pipeline, which would let the code under test decide how far it is trusted.
Effective definition The fully-resolved pipeline definition — templates expanded, digests pinned, parameters bound — recorded immutably per run. What actually ran, as opposed to what the repository says today. The pipeline file at HEAD, which is mutable and therefore useless as a run record.
Sandbox The single-use, hardware-isolated environment created for one job and destroyed after it. The trust boundary; the only place tenant code executes. A 'runner', which in most CI vocabularies is a long-lived host that executes many jobs.
Attestation A signed statement composed by the control-plane attestor naming the commit, definition digest, dependencies and builder for one artefact digest. The non-forgeable link between source and deployed bits. A build-signed statement, which is a claim by the thing being described.
Digest The content hash of an artefact's bytes, and its permanent identity. What is promoted, verified, deployed and rolled back to. A tag or version, which is a mutable pointer to a digest.
Environment A named, governed target holding a pointer to one digest, its configuration, its approvers and its gates. The unit of promotion, locking and rollback. A cluster or namespace, which is where an environment happens to run.
Gate A condition evaluated outside the pipeline before an environment pointer may move. The enforceable control on what reaches production. A gate job inside the pipeline, which the pipeline's author can weaken.
Lease A time-bounded claim a dispatched job must renew. How runner loss is detected, and why a partition resolves to re-dispatch rather than to success. A heartbeat, whose absence is often read as failure or, worse, as completion.
Platform failure A job outcome caused by the platform rather than by the code under test. The class that is retried automatically and counted in the infrastructure-failure ratio. A generic 'error', which conflates infrastructure, cancellation, timeout and policy denial.
Infrastructure-failure ratio Platform failures as a share of all jobs, measured weekly. The single number that decides whether engineers read a red build or reflexively retry it. Build success rate, which mixes the platform's reliability with the code's quality.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.