Architecture Decision Record
Solution Architecture v1.0 · Amazon Web Services · Platform Architecture · 2026-09
CI/CD Platform · Solution Architecture v1.0 · Amazon Web Services · Platform Architecture · 2026-09
The argument these decisions serve is summarised in the Architecture One-Pager.
Seventeen decisions make up this architecture. Everything else across the twenty-two views is either a consequence of one of them or a detail that could be decided differently next quarter without anybody having to redraw the set. Each record carries more than the classic context / decision / consequences triple: the forcing question, how the decision is actually realised on the chosen stack, the alternatives including those that are right for a different organisation, the conditions that would flip the choice, why the choice should still hold as scale and technology change, and the transferable lesson.
Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a product company of 4,200 engineers in 610 teams across 38,000 repositories, running 210,000 pipeline runs and 1.9M jobs a day, peaking at 3,200 jobs a minute with 40,000 concurrent job slots, a median job of 3m 10s and a p95 of 22 minutes, 480 TB of retained artefacts and 6 TB a day of raw log output, with roughly 300 fork pull requests a quarter from outside the company. A reviewer who disagrees with a number can follow it to the decision that depends on it; that is what the numbers are for.
How to read a record
- Question: The forcing question: why a decision was needed at all.
- Context: The requirement, the scale and the constraint that make it hard.
- Decision: What this architecture does, stated so it can be checked.
- How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
- Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- Consequences: What the choice buys and what it costs, both kept visible.
- Choose differently when: The conditions that would flip the decision for your system.
- Why it holds up over time: What keeps the decision right as scale, staff and technology change.
- Lesson: The principle that transfers beyond this platform.
Decision map
Isolation — running code nobody reviewed: The decisions that make it ordinary rather than dangerous to execute a stranger's code on the company's machines.
- ADR-01 · Every sandbox is single-use, created for one job and destroyed after it
- ADR-02 · Isolation is hardware-level, and it does not vary by trust class
- ADR-03 · Trust is a property of the run, assigned at admission from the event
- ADR-04 · The cache is read by everything and written only by trusted runs
- ADR-05 · All sandbox egress is mediated by a policy proxy, on by default
Trust and provenance: The decisions that make the link between commit, build and deployed artefact non-forgeable.
- ADR-06 · The attestor signs provenance; the build never signs anything
- ADR-07 · Credentials are minted per job and revoked when the job ends
- ADR-08 · Attestations are recorded in an append-only log that can be verified without the platform
Fairness and capacity: The decisions that stop one tenant's burst becoming everyone else's queue.
- ADR-09 · Fairness is enforced in the scheduler, by entitlement and weighted fair queuing
- ADR-10 · The warm pool is tenant-agnostic, and a claimed sandbox is never returned to it
- ADR-11 · Interruptible capacity is confined to the background tier
Run state and recovery: The decisions that make a partition resolve to re-dispatch rather than to a false green.
- ADR-12 · The relational store is authoritative for scheduling; the event log is authoritative for history
- ADR-13 · Runner loss is detected by lease expiry, and a job is never successful on an absent runner
- ADR-14 · Six job outcomes, and a platform failure is never reported as the user's
Promotion and gating: The decisions that make "did we ship what we tested" answerable, and the gates unweakenable by the author.
- ADR-15 · Build once, promote the digest, bind configuration at deployment
- ADR-16 · Gates are evaluated outside the pipeline they govern, and fail closed in production
Cost, retention and developer trust: The decisions that keep the safe path the cheap path, and a red build worth reading.
- ADR-17 · Retention is set by artefact provenance class and pipeline class, not globally
Technology by capability
Amazon Web Services was chosen for this exercise deliberately, and for two reasons that pull in the same direction. The first is rotation: across the repository's previous use cases, Azure and self-hosted open-source stacks dominate, and AWS and Google Cloud are the two least leaned on. The second is that the topic genuinely has a cloud-shaped centre of gravity — build isolation at this volume needs nested-virtualisation-capable instances running a lightweight microVM hypervisor, which is where the real hosted CI vendors run and where the primitive originated. The requirement in ask.md stays vendor-neutral throughout; the table below is the architecture's answer, not the requirement's.
| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| Job isolation | Nested-virtualisation-capable EC2 families running a lightweight microVM hypervisor, one VM per job | AWS + open source | gVisor or hardened containers on GKE; Azure Container Instances with Hyper-V isolation | Only hardware virtualisation answers "a build rooted the kernel", and this is the substrate where per-job microVMs are cheap enough for 1.9M jobs a day | ADR-02 |
| Warm pool | Per-host pool of pre-booted, tenant-agnostic microVMs sized from arrival rate | Built | Per-tenant pre-warmed pools with tenant image layers | One fungible pool keeps the idle-waste ceiling a single manageable number across 610 tenants with uneven arrival | ADR-10 |
| Job queue | Amazon SQS with a durable accept-before-acknowledge contract | AWS managed | Kafka or NATS with consumer groups; Cloud Tasks | An accepted run must survive any single component's loss, and the queue is work-to-do rather than the account of what happened | ADR-12 |
| Run metadata | Aurora PostgreSQL, multi-AZ, RPO 0 / RTO 15 min | AWS managed | Spanner; Azure SQL with a failover group | Dispatch decisions are transactional — lease, entitlement, state — and need strong consistency on the hot path | ADR-12 |
| Run event log | Amazon Kinesis, ordered per run, retained 400 days | AWS managed | Kafka with a compacted history topic; Pub/Sub with an archive sink | History, replay and recovery need an ordered immutable account that survives the loss of current state | ADR-12 |
| Artefact store | S3, content-addressed by digest, retention class written at seal time | AWS managed | GCS with object holds; ADLS Gen2 with immutability policies | Immutable bits with a mutable pointer is what makes promote-by-digest and rollback-by-pointer possible | ADR-15 |
| Transparency log and audit | S3 with Object Lock in a separate custody account, 7 years, externally verifiable | AWS managed | A Sigstore-style public log; an append-only ledger service | Evidence held only by the platform it exonerates is not evidence; write-once and externally checkable is the whole requirement | ADR-08 |
| Signing | Hardware-backed KMS key reachable only by the attestor's role | AWS managed | CloudHSM for offline custody; an in-cluster signer with SPIRE identity | The signing key must be unreachable from any isolation host, which is a network and IAM property rather than a cryptographic one | ADR-06 |
| Job identity | OIDC workload identity minted per job, exchanged with STS for short-lived roles | AWS + OIDC | SPIFFE/SPIRE with a workload API; Azure workload identity federation | A leaked credential's useful life should be the job's life, and its authority the job's authority | ADR-07 |
| Control plane runtime | EKS across three AZs, scheduler sharded with a leader per shard | AWS managed | ECS on Fargate; GKE Autopilot | Admission, gating and attestation must scale independently, and the scheduler needs cross-tenant state that shards rather than replicates | ADR-09 |
| Cache and mirror | S3-backed content-keyed cache, read-wide and write-privileged, plus a regional dependency mirror | AWS + built | A registry-backed cache; a self-hosted Artifactory or Nexus mirror | Splitting the write privilege removes the most-exploited CI vulnerability class, and the mirror is what makes restrictive egress fast enough to keep | ADR-04 |
| Egress control | No default route from sandboxes; a policy-enforcing proxy with per-tenant allow-lists | Built | Cloud-native egress firewall rules; a service mesh egress gateway | The restrictive posture has to be the default one, and denial has to read as a policy decision rather than a network fault | ADR-05 |
| Gate decisions | Policy decision service in the control plane, versioned bundles cached per tenant | Built + open-source policy engine | In-pipeline gate jobs; admission policy in the target platform | A pipeline author must not be able to weaken the gates applying to their own deployment | ADR-16 |
| Interruptible capacity | Spot instances for the background tier only, drain notice honoured, re-dispatch to on-demand | AWS managed | Spot for everything with transparent retry; on-demand only | Cheap capacity is paid for in tail latency, so it belongs where nobody is waiting at the end of the job | ADR-11 |
| Logs and metrics | Agent-streamed logs tiered from searchable to archival on S3, class-based retention | AWS + built | OpenSearch for the whole window; a managed log product at full retention | 6 TB a day is the fastest-growing cost line, and retention by pipeline class is the only way to spend it where it is wanted | ADR-17 |
The decisions, and the alternatives that lost
Isolation — running code nobody reviewed
The decisions that make it ordinary rather than dangerous to execute a stranger's code on the company's machines.
ADR-01 · Every sandbox is single-use, created for one job and destroyed after it
Status: Accepted · Shown on views: 04, 13, 20
Does a build environment get reused between jobs, and if so, what exactly is guaranteed to have been removed?
Context. Reusing a warm build environment is the single largest performance win available to a CI platform. A pool of containers with the dependency layers already pulled and the toolchain already warm turns a forty-second start into a two-second one, and every CI vendor has been tempted by it. The problem is that the guarantee required is not "we cleaned up the working directory". It is "nothing the previous job wrote — a file outside the workspace, a kernel object, a cached credential in an agent socket, a modified binary on the PATH, an entry in a package manager's global store — can influence or be read by this one". That guarantee is impossible to audit and its violations do not fail tests; they fail silently, in one job in ten thousand, and look like a flake.
Decision. A sandbox is created for exactly one job and destroyed after it. No sandbox serves two jobs, and the destruction happens whether the job succeeded, failed, timed out or was killed. Warmth is bought back by pre-booting sandboxes in a pool before they are claimed, never by reusing one after a job has run inside it.
How it is realised on AWS. The sandbox manager maintains a warm pool of pre-booted microVMs on each isolation host. A microVM in the pool has booted a base image and has never executed tenant code; at dispatch, one is claimed, the job's workspace and credentials are injected, and the VM is marked non-reusable. On completion the VM is terminated and its backing storage discarded, and the pool replenishes asynchronously from the host's capacity. A VM that has been claimed is never returned to the pool, even if the job never started.
| Option | Verdict | Reasoning |
|---|---|---|
| Single-use sandbox per job, warm pool of never-used sandboxes | Chosen | Makes residue impossible rather than unlikely, and moves the performance problem to cold start where it is measurable |
| Reused container pool with a cleanup step between jobs | Rejected | Cheapest and fastest; the cleanup guarantee cannot be audited and its failures look like flakes |
| Reused for trusted builds, single-use for fork builds | Rejected | Assumes a trusted build cannot be hostile, which a compromised dependency disproves |
| A whole dedicated instance per job | Right elsewhere | Right where jobs are long and few, or where regulatory isolation is per-customer; the cold start and cost do not survive 1.9M jobs a day |
What it buys
- Cross-job contamination stops being a class of bug, which removes an entire category of unreproducible flake
- The teardown path is exercised 1.9M times a day, so it is reliable by the time it matters
- Capacity accounting becomes simple: one job occupies one sandbox for a measurable duration, which is what makes cost per job-minute meaningful
What it costs
- Cold start becomes the platform's hardest performance problem and the warm pool a permanent cost line — up to 8% of compute minutes idle by design
- Dependency warmth must be recovered through the cache and mirror rather than through a reused filesystem, which makes cache hit rate load-bearing
- Cost per job-minute is materially higher than a reused-container platform's, and that gap is visible to anyone comparing vendors
Choose differently when. If every repository in the estate were internal, every contributor employed and vetted, and every dependency vendored and reviewed before entry, the residue risk would be low enough that a reused pool with a cleanup step would be the better economic choice. The decision is justified by roughly 300 fork pull requests a quarter and by a dependency graph the company does not control — remove both and the calculation changes.
Why it holds up over time. The rule is about lifecycle, not about technology. Whatever the isolation primitive becomes — a lighter hypervisor, a hardware partition, something not yet built — "one job, then destroyed" remains expressible and remains the property that makes every other isolation guarantee meaningful. The cost of the rule falls as boot times fall; the cost of abandoning it does not fall at all.
Lesson. When a guarantee cannot be audited, do not try to enforce it with a cleanup step. Change the lifecycle so the guarantee is structural, and pay for it somewhere you can measure.
ADR-02 · Isolation is hardware-level, and it does not vary by trust class
Status: Accepted · Shown on views: 15, 16, 20
Must the boundary between two concurrently running jobs survive a kernel exploit, and does the answer differ for a build of the company's own code?
Context. Containers isolate by kernel namespace, which means a kernel vulnerability reachable from inside the container is a path out of it. Syscall filtering and user-space kernels narrow that surface considerably but do not remove the shared-kernel premise. The question is not whether such exploits are common — they are rare — but what the consequence is when one lands: on a shared-kernel host, another tenant's job, its credentials and its source are in reach. The second half of the question is more interesting. It is tempting to run fork builds in microVMs and trusted builds in containers, because most jobs are trusted and containers are three times cheaper. That reasoning assumes a trusted build cannot be hostile, which has not been true since dependency compromise became the dominant supply-chain attack: a malicious postinstall script in a transitive dependency of the company's own service executes with exactly the privileges of a trusted build.
Decision. Every job runs in a hardware-virtualised microVM. Isolation strength is identical for a trusted branch build, a fork build, a scheduled run and an interactive session. What varies between trust classes is what the sandbox is given — credentials, cache write authority, publish identity — never how well it is contained.
How it is realised on AWS. Nested-virtualisation-capable EC2 instance families run a lightweight microVM hypervisor, one VM per job, with the per-job agent inside the VM and the sandbox manager outside it. Jobs that declare nested container execution get it inside their own VM without host privileges. Two instance families are qualified at all times so a capacity shortage in one is not a platform-wide throughput event.
| Option | Verdict | Reasoning |
|---|---|---|
| microVM per job for every trust class | Chosen | One substrate to harden, one cold-start number to optimise, and no capacity stranded in the wrong pool |
| Hardened containers with syscall filtering for all jobs | Rejected | Fastest and cheapest, and adequate until the first kernel escape, at which point it is adequate for nothing |
| microVMs for untrusted, containers for trusted | Rejected | The common design; it prices the fork risk correctly and the dependency-compromise risk at zero |
| Separate physical fleets per trust class | Deferred | Worth revisiting if a regulator requires physical separation; it doubles the capacity-planning problem for a marginal gain over microVMs |
What it buys
- A kernel-level compromise inside one job yields no access to another, which is the only honest answer to "what if a build roots the kernel"
- One substrate means one hardening effort, one cold-start budget and one capacity pool — no stranded capacity in a fleet nobody is using this hour
- Nested container execution can be offered safely, because the nesting happens inside the tenant's own VM
What it costs
- Higher cost per job-minute and a slower start than containers, borne on every job rather than only on risky ones
- Dependence on a narrower instance market, which is a capacity-planning risk with no software mitigation
- Some workloads that assume host features will need explicit support rather than working by accident
Choose differently when. If the estate had no external contributors and every dependency were vendored and reviewed, the dominant residual risk would be a platform bug rather than tenant code, and hardened containers would be the right economic answer. Equally, if a future container runtime offered a genuine hardware-backed boundary at container cost, this decision becomes an implementation detail rather than a trade-off.
Why it holds up over time. The requirement is "the boundary survives a kernel exploit", which is a property, not a product. It has been satisfied by different technologies each decade and will be satisfied by others. Stating it as a property rather than as "we use microVMs" is what lets the substrate be replaced without revisiting anything else in this record set.
Lesson. Do not price a risk class by how often you expect its trigger. Price it by what is reachable when the trigger fires, and then refuse to grade the boundary by how much you trust the code — because the code you trust is exactly what an attacker wants to arrive inside.
ADR-03 · Trust is a property of the run, assigned at admission from the event
Status: Accepted · Shown on views: 05, 12, 15
At what moment, and from what evidence, does the platform decide whether this run's code is trusted — and can that decision ever be revised upward?
Context. Every credential and cache decision in the platform depends on one bit: is the code about to execute code the company has reviewed? Getting that bit from the wrong place is the root of most CI security incidents. If it is derived from the pipeline definition, the fork's own definition can claim to be trusted. If it is derived at credential-request time, a job that has already started can escalate. If it is recomputed per step, a step can change it. And if it is computed from the pull request's target repository rather than from the head's provenance, a pull-request trigger that runs the head's code with the base's permissions is exactly the well-documented mistake that has leaked deployment credentials from several widely-used CI systems.
Decision. The trust class is computed once, at admission, from the trigger event and the repository's membership records — never from anything inside the pipeline definition or the commit. It is written as an immutable column on the run and attached to every job, credential request and cache operation for the run's life. No path exists to raise it. Lowering it is possible only by cancelling the run.
How it is realised on AWS. The admission service reads the signed webhook payload, resolves the head repository and contributor against the tenant's membership and approval records, and writes trust_class on the run row before the run enters the queue. The job identity minted by identity federation binds tenant, repository, ref and trust class into the token, so the secret broker and cache service evaluate the class from the token rather than by looking it up again. A run's effective definition is recorded alongside the class, so an auditor can see both what ran and under what posture.
| Option | Verdict | Reasoning |
|---|---|---|
| Computed once at admission from the event, immutable on the run | Chosen | One place to get right, one column to audit, and no escalation path to reason about |
| Evaluated per credential request against current state | Rejected | Flexible, and creates a window in which a membership change mid-run changes a running job's privileges |
| Declared in the pipeline definition with policy validation | Rejected | Lets the artefact under test participate in deciding how much it is trusted |
| Maintainer-approved promotion of a fork run to trusted | Right elsewhere | Reasonable for a small open-source project where one person reviews every diff; at 610 tenants it becomes a click that nobody reads |
What it buys
- The security posture of a run is a single auditable value, set before any capacity is committed
- Credential and cache services need no membership lookups on the hot path, because the class travels in the token
- The fork case becomes ordinary: it is a class, not an exception handled by a special code path
What it costs
- A contributor who becomes trusted mid-run does not benefit until the next run, which occasionally confuses maintainers
- The classifier is a small piece of code with disproportionate consequence, and needs test coverage out of proportion to its size
- A legitimate need to run a fork's code with credentials has no in-platform answer, only the trusted re-run at merge
Choose differently when. If the platform served a single team with no external contributors, the class would always be the same value and the machinery would be pure overhead. And if a future source-control platform offered a cryptographically attested contributor-trust signal at event time, the classifier would become a thin adapter over it rather than a decision the platform makes itself.
Why it holds up over time. "Compute the security-relevant fact once, from the least manipulable input available, and make it immutable" is a design rule older than CI and will outlast this platform. What changes over time is which input is least manipulable; the rule does not.
Lesson. Never let the thing being evaluated contribute to the evaluation. Derive the security-relevant fact from the event, write it down, and give yourself no code path that can raise it later.
ADR-04 · The cache is read by everything and written only by trusted runs
Status: Accepted · Shown on views: 05, 14, 20
Who is allowed to write a cache entry that a later build will restore and execute?
Context. A build cache is not passive data. Restored into a workspace, it becomes compiler output, dependency binaries, and sometimes executables on the PATH. If an untrusted run can write an entry that a trusted run later restores, the untrusted run has achieved code execution inside the trusted build — and therefore inside whatever the trusted build is allowed to publish and deploy. This is not hypothetical: it is the most-exploited class of CI vulnerability in practice, because the cache is the one shared mutable surface a fork build legitimately touches. The counter-pressure is real too. Cache hit rate is the single largest determinant of build latency, and the contributors most in need of a fast build are exactly the ones who cannot be allowed to write.
Decision. Any run may read the cache. Only a run classified as trusted may write it. Every entry is integrity-verified on restore, and a verification failure is treated as a cache miss rather than as an error. The slower fork build is accepted as the price, and the platform tells the contributor why.
How it is realised on AWS. Cache write credentials are issued by the broker only against a job identity carrying the trusted class; an untrusted job's request is refused with a distinct, explicit reason rather than silently returning nothing. Entries are content-keyed with declared fallback keys and carry a digest verified on restore. Restore is served from a read-wide path that requires no write credential at all, so a fork build's cache read is not a privilege it could misuse.
| Option | Verdict | Reasoning |
|---|---|---|
| Read wide, write privileged, integrity-verified on restore | Chosen | Keeps most of the hit rate and removes the code-execution path entirely |
| Shared per-repository cache writable by any branch or fork | Rejected | Best hit rate, and a direct code-execution path from a stranger into the production build |
| Per-branch caches with read fallback to the default branch | Rejected | Better than shared-writable, and still lets an untrusted branch's entry be read by a trusted build via fallback |
| Content-addressed cache with attested writers only | Deferred | The right long-term answer and a natural extension of the attestor; deferred because it needs the remote-execution work in Phase 3 |
What it buys
- The most-exploited CI vulnerability class is removed by credential design rather than by validation
- A poisoned or corrupt entry costs a slow build, never a compromised one, because verification failure is a miss
- Trusted builds keep essentially the full benefit of the cache, which is where the majority of the job volume is
What it costs
- Fork builds are measurably slower, and the contributors affected are the ones the company least wants to frustrate
- A fork build that would have warmed the cache for its own next iteration cannot, so iteration on a fork PR is slow throughout
- Cache hit rate must be reported per trust class, or the aggregate number hides the fork experience
Choose differently when. If the cache held only content-addressed, independently verifiable artefacts whose provenance was attested — so that restoring an entry proved who produced it and from what — then write access would stop being a privilege and any run could write safely. That is the deferred option above, and it is the direction this decision should eventually move in.
Why it holds up over time. The asymmetry between reading shared state and writing it is permanent, and it gets more valuable as caching gets more aggressive. Remote execution and distributed action caches raise the stakes rather than lowering them: the more of a build comes from cache, the more a cache write is worth to an attacker.
Lesson. Shared mutable state between security domains is a code-execution path, whatever it is called. Split the privilege before optimising the hit rate, and be honest with the people the split slows down.
ADR-05 · All sandbox egress is mediated by a policy proxy, on by default
Status: Accepted · Shown on views: 15, 20, 22
What can a running job reach on the network, and who decides?
Context. A job needs the network: it fetches dependencies, pulls base images, clones submodules, sometimes calls a test double. Unrestricted egress makes all of that work and also makes exfiltration trivial — a build that has read something it should not can simply post it somewhere. Unrestricted egress is also how a build defeats its own recorded dependency set: a step that curls a script at build time makes the resolved-dependency record a fiction. The difficulty is that a strict allow-list breaks builds in ways that are tedious to diagnose, so any design that makes restriction opt-in results in nobody opting in.
Decision. Every sandbox's egress passes through a policy-enforcing proxy. Restriction is on by default: an untrusted run reaches only the dependency mirror and source control. A tenant may define an allow-list for its trusted runs. Every allowed and denied destination is logged and attributable to the run.
How it is realised on AWS. Isolation hosts place sandboxes on a network with no default route; the per-job agent's traffic is directed to a proxy that evaluates destination against the tenant's policy and the run's trust class. Package and image pulls are served from the platform's own mirror, which is both the performance path and the audit path. Denials return a distinct, named error so the failure reads as a policy decision rather than as a network fault.
| Option | Verdict | Reasoning |
|---|---|---|
| Mediated egress, restrictive by default, tenant allow-list for trusted runs | Chosen | The safe posture is the one you get without doing anything, and the mirror makes it fast rather than merely safe |
| Open egress with monitoring and alerting | Rejected | Breaks nothing and prevents nothing; the alert arrives after the exfiltration |
| Opt-in allow-list per pipeline | Rejected | Correct in principle; in practice almost no pipeline opts in, so the control does not exist |
| No egress at all, everything pre-provisioned into the sandbox | Right elsewhere | Right for a classified or air-gapped environment; the provisioning burden is not worth it when a mirror will do |
What it buys
- Exfiltration from a hostile build requires defeating the proxy, not merely running code
- The recorded dependency set becomes trustworthy, because fetching outside it is visible and usually blocked
- Upstream registry outages degrade build latency rather than stopping all builds, because the mirror is the default path
What it costs
- Builds that reached the internet incidentally now fail, and the first month of adoption is spent adding legitimate destinations
- The mirror becomes a critical dependency with its own availability and storage cost
- A determined build can still tunnel over an allowed destination; the control is a barrier, not a proof
Choose differently when. If the platform served only internal code with no untrusted runs and no supply-chain concern, the diagnostic cost of the allow-list would outweigh its benefit and monitored open egress would be defensible. As soon as one fork build runs, it is not.
Why it holds up over time. Mediating egress is independent of the network technology that implements it, and its value rises with every step the industry takes towards supply-chain attacks. The mirror that makes it palatable is also the thing that makes build reproducibility possible, so the investment pays twice.
Lesson. A control that must be turned on will not be. Make the restrictive posture the default and invest in making it fast, because the diagnostic pain — not the security argument — is what decides whether the control survives contact with a deadline.
Trust and provenance
The decisions that make the link between commit, build and deployed artefact non-forgeable.
ADR-06 · The attestor signs provenance; the build never signs anything
Status: Accepted · Shown on views: 02, 13, 20
Who makes the statement "this artefact was built from this commit by this pipeline", and why should anyone believe it?
Context. A provenance attestation is only worth the difficulty of forging it. If the signing key is available inside the build — as an environment variable, a mounted file, or a credential the job can exchange — then the statement is made by the thing being described, and a compromised build can produce an artefact together with a perfectly valid attestation saying it came from somewhere else. That is not a marginal weakness; it makes the entire evidence chain decorative, because the case the chain exists to detect is precisely a compromised build. The counter-argument is practical: the build knows most of the facts, so having it assemble and sign the statement is much simpler than reconstructing them outside.
Decision. The signing key is held by a control-plane attestor and is never present inside a sandbox. The job reports its outcome and uploads its artefact; the attestor composes the attestation from the inputs the platform itself observed — the resolved effective definition, the source commit, the recorded dependency set, the builder identity and the timings — and signs it. The build's own claims are inputs to be recorded, never assertions to be signed.
How it is realised on AWS. The signing key lives in a hardware-backed key store reachable only by the attestor's role; no path exists from an isolation host to it. The attestor is invoked by the sandbox manager after the job reports, binds the request to the job's lease and identity, composes the statement from control-plane records, signs, and appends to the transparency log. The artefact is addressed by the digest of its bytes, so the statement is about specific bits rather than about a name.
| Option | Verdict | Reasoning |
|---|---|---|
| Control-plane attestor signs from observed inputs | Chosen | A compromised build can produce a bad artefact but not a credible claim about one |
| The build signs with a short-lived key issued to the job | Rejected | Much simpler, and the statement is then made by the code under suspicion |
| The build composes the statement, the control plane signs it unexamined | Rejected | Looks like the chosen option and is the rejected one: the content is still the build's |
| A separate rebuilder independently reproduces and signs | Deferred | The strongest available evidence and roughly doubles build cost; worth it for a small set of release artefacts, not for 1.9M jobs a day |
What it buys
- An attestation is worth verifying, which is the entire point of producing one
- Verification can be performed outside the platform against the transparency log, so the platform is not the only witness to its own claims
- The gate service has something objective to check before a deployment, rather than a self-report
What it costs
- The attestor must reconstruct facts the build knew directly, which means the control plane has to record more than it otherwise would
- The attestor becomes the highest-consequence code in the platform: a confused-deputy bug there yields valid, false statements
- Signing is on the critical path of every job completion, adding latency and a dependency to the seal step
Choose differently when. If artefacts were never deployed anywhere consequential — a research estate, a throwaway environment — the cost of a control-plane attestor would exceed its value and a build-signed statement would be adequate record-keeping. The decision is justified by the artefact reaching production under a gate that trusts the statement.
Why it holds up over time. "The observer signs, not the observed" predates software supply chains and will outlive every signing format. Sigstore, in-toto, whatever replaces them: each is a way of expressing this separation, and the separation is the part worth committing to.
Lesson. Evidence produced by the subject of the evidence is not evidence. If a claim must survive the compromise of the thing it describes, the claim has to be made somewhere else.
ADR-07 · Credentials are minted per job and revoked when the job ends
Status: Accepted · Shown on views: 13, 20, 21
What credential is inside a running job, how long is it useful, and what happens to it when the job stops?
Context. A long-lived credential inside a build environment is the most valuable thing an attacker can obtain from a CI platform, because it outlives the build, works from anywhere, and usually has more authority than the build needed. Static secrets in a pipeline's configuration are the classic form; a cloud role attached to the runner host is the modern one, and it is worse, because every job on that host inherits it whether it needed it or not. The alternative — minting an identity per job and exchanging it for narrowly scoped, short-lived credentials — puts identity federation on the hot path of every one of 1.9M daily jobs, which is a real availability and latency cost.
Decision. Every job is issued a workload identity binding tenant, repository, ref and trust class, valid for the job's lifetime. All secret and cloud access is scoped to that identity. Nothing long-lived is present in a sandbox. Revocation is pushed at job end rather than left to token expiry, and every secret read is recorded in the audit trail before the value is returned.
How it is realised on AWS. The scheduler asks identity federation to mint the job identity at dispatch; the token is delivered to the per-job agent and never to the job's own environment except as an explicitly requested, scoped exchange. The secret broker evaluates the token's trust class against policy, writes the audit record, then returns the value with redaction registered on the log path. Cloud access is obtained by exchanging the identity with the cloud's token service for a short-lived role. At job end the scheduler revokes the identity rather than waiting for the TTL.
| Option | Verdict | Reasoning |
|---|---|---|
| Per-job identity, scoped exchange, revoked at job end | Chosen | A leaked credential's useful life is the job's life, and its authority is the job's authority |
| Host-attached instance role shared by every job on the host | Rejected | Simplest and requires no federation; every job inherits authority it did not ask for |
| Static secrets in pipeline configuration | Rejected | Still the industry's most common design and the source of most credential leaks |
| Per-step identity with per-step scope | Deferred | Strictly better and multiplies federation traffic by the step count; revisit if step-level blast radius becomes the binding concern |
What it buys
- A credential exfiltrated from a build is worthless minutes later and was never broadly scoped
- Every secret read has an audit record naming the job, the repository, the ref and the trust class
- Denying an untrusted run is a policy evaluation on a token rather than a special case in the build path
What it costs
- Identity federation is on the hot path of every job and must be more available than the execution plane it serves
- Token minting and exchange add latency at job start, inside an already-tight cold-start budget
- Pipelines that expect a static secret must be migrated, and the migration is visible to every team
Choose differently when. If the platform's jobs needed no external authority at all — pure compile-and-test with no publish and no cloud access — the federation machinery would be unnecessary and a host role with no permissions would do. The moment a job can publish or deploy, it does not.
Why it holds up over time. Short-lived, narrowly scoped, workload-bound credentials are the direction every identity system has moved in for fifteen years and the direction the cloud providers keep extending. Building on federation rather than on stored secrets means the platform inherits those improvements instead of having to migrate to them.
Lesson. Measure a credential by how long it is useful after it leaks and by how much it can do that its holder did not need. Both numbers should be as close to zero as the platform's latency budget allows.
ADR-08 · Attestations are recorded in an append-only log that can be verified without the platform
Status: Accepted · Shown on views: 11, 20, 21
If the platform itself were compromised or simply wrong, how would anybody find out that an attestation had been altered or invented?
Context. An attestation stored in a mutable database beside the run metadata is trustworthy exactly as far as the platform is trustworthy. That is a circular guarantee, and it fails in the two cases evidence is most needed: a compromise of the platform's own control plane, and an honest bug that mis-attributes one job's inputs to another's artefact. Both produce records that are internally consistent and factually wrong, and neither is detectable from inside. The cost of doing better is real: an append-only, tamper-evident log is more operationally awkward than a table, cannot be corrected in place, and forces the platform to keep records it might prefer to tidy.
Decision. Every attestation is appended to a tamper-evident, append-only log, retained seven years, verifiable by a third party without the platform's cooperation. Nothing in the log is ever edited or deleted. An artefact whose provenance fails verification is quarantined rather than removed, because a deleted artefact cannot be investigated.
How it is realised on AWS. Attestations and SBOMs are written to object storage under an immutability policy, with the log's inclusion proofs published so an external verifier can check an entry without asking the platform to vouch for it. The transparency log sits in the custody zone, reachable from the attestor and from nothing in the execution plane. Cross-region replication carries evidence at RPO 0, because losing proof is worse than losing work.
| Option | Verdict | Reasoning |
|---|---|---|
| Append-only tamper-evident log, externally verifiable, 7 years | Chosen | Makes an inconsistency discoverable without the platform's cooperation, which is the only case that matters |
| Attestations as rows beside run metadata | Rejected | Simple, queryable, and worthless in the two failure cases the evidence exists for |
| Public transparency log shared with the wider ecosystem | Right elsewhere | Right for open-source artefacts the world consumes; for internal artefacts it leaks the shape of the estate |
| Periodic signed snapshots of the attestation table | Rejected | Cheaper, and detects tampering only at snapshot granularity, which an attacker chooses to work inside |
What it buys
- An auditor can answer "which commit produced this image" without the platform being a trusted intermediary
- A control-plane bug that mis-signs becomes discoverable after the fact rather than invisible
- Quarantine rather than deletion keeps the investigation possible, and the artefact still cannot deploy
What it costs
- Nothing can be corrected in place; an error is fixed by appending a correction, which is operationally unfamiliar
- Seven years of immutable evidence is a storage and legal-hold commitment that grows monotonically
- The immutability policy also protects records the company might later wish to remove, which is the point and is sometimes inconvenient
Choose differently when. If artefacts were short-lived and never subject to audit, retention of any kind would be waste and a table would suffice. The decision is driven by two audit reviews a year and by production images whose origin must be provable years after the engineer who built them has left.
Why it holds up over time. Tamper-evident append-only logs have outlived several generations of signing technology, and the reason is structural: the guarantee comes from the data structure rather than from the operator. Whatever replaces today's log formats will still be a log, and its entries will still need to be verifiable by someone who does not trust the writer.
Lesson. Evidence held only by the party it exonerates is not evidence. Make the record append-only and externally checkable, and accept that you can no longer tidy it.
Fairness and capacity
The decisions that stop one tenant's burst becoming everyone else's queue.
ADR-09 · Fairness is enforced in the scheduler, by entitlement and weighted fair queuing
Status: Accepted · Shown on views: 02, 03, 18
When 40,000 slots are full and one tenant has just pushed a 900-job monorepo pipeline, whose job runs next?
Context. A global FIFO queue with per-tenant concurrency caps is the usual first design and it is comprehensible, cheap and wrong in one specific way: a cap bounds how much capacity a tenant holds but not how much of the queue it occupies. A monorepo pipeline that admits 900 jobs at once puts 900 entries ahead of the next tenant's five, so the small tenant's wait is set by the large tenant's burst regardless of the cap. Because CI latency is felt personally — an engineer is sitting there — the small tenant experiences this as the platform being broken, and the platform team experiences it as an email asking large teams to be considerate. Weighted fair queuing fixes it, at the cost of needing cross-tenant state in the dispatch decision.
Decision. Each tenant has a concurrency entitlement and a priority weight. Dispatch selects across tenants by weighted fair queuing rather than by arrival order, so queue wait inside an entitlement is bounded independently of any other tenant's burst. Idle entitlement is lent to tenants that are over theirs, and reclaimed when the owner returns. Queue wait inside entitlement is an SLO; a tenant over its entitlement has no wait guarantee.
How it is realised on AWS. The scheduler is sharded, with per-shard fair-queue state and a leader per shard, and entitlements held in the run metadata store. Priority tiers — interactive, scheduled, background — are evaluated before fairness, so an interactive job outranks a background job of any tenant, and background work is preemptible. Queue position and estimated wait are computed from the same state and exposed to the waiting engineer.
| Option | Verdict | Reasoning |
|---|---|---|
| Per-tenant entitlement with weighted fair queuing and lending | Chosen | Bounds the small tenant's wait and still lets a big tenant use idle capacity |
| Global FIFO with per-tenant concurrency caps | Rejected | Simple and cheap; the small tenant's wait is set by the large tenant's burst |
| Hard per-tenant partitions, no lending | Rejected | Perfectly fair and strands capacity inside unused entitlements, which at 610 tenants is most of the fleet most of the time |
| A market with priced priority | Deferred | Aligns incentives elegantly and requires a chargeback culture the organisation does not yet have; revisit alongside showback |
What it buys
- A two-person team's five-job pipeline has a bounded wait no matter what the monorepo is doing
- Idle entitlement is usable, so fairness does not cost utilisation
- Queue position and estimated wait become computable, which turns an unexplained wait into an explained one
What it costs
- The scheduler needs cross-tenant state in the dispatch decision, which constrains how far it can be sharded
- Entitlements become a thing platform engineers must curate, and a badly-set entitlement is now the cause of a wait
- Fairness violations need their own monitoring, because the failure is silent from any single tenant's point of view
Choose differently when. If tenant sizes were roughly uniform and bursts rare, FIFO with caps would deliver nearly the same outcome for a fraction of the complexity. The decision is justified by a distribution in which a handful of monorepos generate a large share of job volume — and that distribution is getting more extreme, not less.
Why it holds up over time. Weighted fair queuing is fifty-year-old theory from packet scheduling and has survived every change of substrate since. The pressure it addresses — the ratio between the largest and smallest workload sharing a resource — increases with monorepo size, generated code volume and AI-assisted development, so the decision gets stronger with time.
Lesson. A concurrency cap limits what a tenant holds, not what a tenant blocks. If wait time is the experience, fairness has to be a property of the queue, not of the quota.
ADR-10 · The warm pool is tenant-agnostic, and a claimed sandbox is never returned to it
Status: Accepted · Shown on views: 13, 16, 18
Can a sandbox be pre-warmed with a tenant's image and dependency layers, and if so, may it later be handed to a different tenant?
Context. Single-use sandboxes (ADR-01) make cold start the platform's hardest performance problem, and the obvious mitigation is to pre-warm with tenant content: pull the base image, hydrate the dependency layers, then hand the ready sandbox to that tenant's next job. The saving is large — most of the queued-to-first-step budget is image pull and dependency hydration. The problem is that a pre-warmed sandbox is now tenant-specific, so it can only serve that tenant; at 610 tenants with uneven arrival rates, most pre-warmed sandboxes sit idle waiting for a job that does not come, and the ones that are needed are in the wrong pool. Handing a tenant-warmed sandbox to another tenant is the alternative, and it reintroduces exactly the residue problem ADR-01 exists to remove.
Decision. The warm pool holds tenant-agnostic sandboxes that have booted a base image and have never seen tenant content. One is claimed at dispatch, the workspace and credentials are injected, and it is never returned to the pool. Tenant-specific warmth is recovered through the cache and the dependency mirror, not through a pre-warmed sandbox.
How it is realised on AWS. Each isolation host maintains a small pool of pre-booted microVMs from a common base image, sized from the observed arrival rate for that host's resource class. Dispatch claims one and marks it non-reusable. Tenant warmth comes from the cache restore path and from the mirror serving image and package layers from the same region, which is where the remaining cold-start budget is spent and optimised.
| Option | Verdict | Reasoning |
|---|---|---|
| Tenant-agnostic pre-booted pool, claimed at dispatch, never returned | Chosen | One pool serves every tenant, so the idle-waste ceiling is a single number to manage |
| Per-tenant pre-warmed pools with tenant image layers | Rejected | Best cold start; at 610 tenants most warm capacity is idle in the wrong pool |
| Tenant-warmed pools that may be reassigned between tenants | Rejected | Has the cold-start benefit and reintroduces the residue problem ADR-01 removes |
| Per-tenant pools for the largest tenants only, agnostic for the rest | Deferred | Defensible hybrid worth revisiting once per-tenant arrival rates are measured rather than assumed |
What it buys
- Idle warm capacity is one fungible pool, so the 8% waste ceiling is a number that can actually be held
- No sandbox ever holds one tenant's content while waiting for another tenant's job, so ADR-01 survives intact
- The cache and mirror become the levers for cold start, and both benefit every tenant rather than the one they were warmed for
What it costs
- Cold start is longer than a tenant-warmed pool would give, and that difference is the most visible performance number the platform publishes
- Cache hit rate and mirror locality carry more weight than they otherwise would, which makes them availability-critical
- A very large tenant with unusual images gets no special treatment, and will ask for it
Choose differently when. If the estate were a handful of large tenants with steady arrival rates and huge images, per-tenant pools would be both affordable and materially faster, and the hybrid above becomes the right answer. The decision is driven by 610 tenants with uneven arrival.
Why it holds up over time. The trade — fungible capacity against specific warmth — is a permanent one in any pooled system, and the right side of it is decided by the number of tenants and the variance of their arrival. Both of those are measurable, so this decision is designed to be revisited with data rather than with argument.
Lesson. Pre-warming with tenant-specific content converts a capacity problem into a placement problem. Before paying for the warmth, check whether the pool will be in the right place when the job arrives.
ADR-11 · Interruptible capacity is confined to the background tier
Status: Accepted · Shown on views: 16, 18, 22
Which jobs may run on capacity the cloud can take back, and what happens when it does?
Context. Interruptible capacity is substantially cheaper and CI is a near-ideal workload for it: jobs are short, stateless and retryable. The temptation is to put everything on it and re-dispatch on reclamation. That works arithmetically and fails experientially, because a reclamation adds the full cold start plus the elapsed work to a job someone is sitting and waiting for, and the distribution of that added latency is exactly the long tail the queue-wait SLO exists to bound. Meanwhile the nightly and backfill work, which nobody is waiting for, is the natural home for the risk.
Decision. Interruptible capacity serves the background tier only: scheduled runs, nightly work, backfills. Interactive and scheduled-but-blocking work runs on on-demand capacity. A reclaimed job is re-dispatched to on-demand capacity, and the drain notice is honoured where the provider gives one. The target is that at least 55% of eligible background job-minutes run on interruptible capacity.
How it is realised on AWS. Resource classes carry an eligibility flag; the scheduler places eligible background jobs on interruptible hosts and everything else on on-demand. The per-job agent handles the drain notice by reporting a distinct platform-failure cause and terminating cleanly, so the re-dispatch is attributable and does not pollute the infrastructure-failure ratio's interpretation. Background work is also preemptible by interactive work on shared hosts.
| Option | Verdict | Reasoning |
|---|---|---|
| Interruptible for background tier only, re-dispatch to on-demand | Chosen | Takes most of the saving without putting the waiting engineer's tail latency at risk |
| Interruptible for everything with transparent re-dispatch | Rejected | Best unit cost and worst p99 for exactly the jobs someone is watching |
| On-demand only | Rejected | Simplest capacity story and leaves a large, easily-harvested saving on the table |
| Interruptible with checkpoint and resume mid-job | Right elsewhere | Right for long ML training jobs; a 3-minute CI job is cheaper to restart than to checkpoint |
What it buys
- A meaningful share of compute cost is harvested from work nobody is waiting for
- The interactive queue-wait SLO is not exposed to reclamation tail latency
- Reclamation becomes an ordinary, attributable event rather than an incident
What it costs
- Nightly throughput is variable, so a capacity-scarce night lengthens the nightly window
- Two capacity pools to plan and monitor rather than one
- The eligibility flag is a decision someone has to make per resource class, and a wrong one is felt as latency
Choose differently when. If the interactive volume were small and latency-insensitive — a batch-oriented estate, or one where CI results are read the next morning — putting everything on interruptible capacity would be correct. Equally, if reclamation rates in the chosen families fell far enough, the distinction would stop earning its complexity.
Why it holds up over time. The rule generalises past any particular cloud's interruptible product: place cheap-but-unreliable capacity where the latency distribution does not matter. That mapping between workload patience and capacity reliability is permanent even as the products change names.
Lesson. Cheap capacity is not free capacity; it is paid for in tail latency. Spend that tail on work with no one waiting at the end of it.
Run state and recovery
The decisions that make a partition resolve to re-dispatch rather than to a false green.
ADR-12 · The relational store is authoritative for scheduling; the event log is authoritative for history
Status: Accepted · Shown on views: 11, 12, 19
When the control plane restarts mid-run, what does it read to find out what was happening — and if two sources disagree, which one wins?
Context. Run state has two consumers with incompatible needs. Scheduling needs a strongly consistent, transactional view it can make decisions against right now: is this job leased, is this tenant at its entitlement, may this run proceed. History, recovery and audit need an ordered, immutable account of everything that happened, including the transitions a current-state row has already overwritten. Serving both from one store means either a relational table that loses history, or an event log that cannot answer "how many slots is this tenant holding" without a fold over everything. Serving them from two stores means they can disagree, and a design that does not say which one wins will discover the answer during an incident.
Decision. Run and job current state live in a transactional relational store, which is authoritative for every scheduling decision. Every transition is also appended to an ordered event log, which is authoritative for history, replay and audit. On disagreement, the relational store wins for scheduling and the event log wins for what happened. The log is retained independently, so losing the metadata store does not lose the account.
How it is realised on AWS. Aurora PostgreSQL holds runs, jobs, leases and entitlements with RPO 0 and RTO 15 min; transitions are appended to a Kinesis stream with RPO 0 and RTO 30 min, retained 400 days. The scheduler reads and writes the relational store transactionally and emits to the log as part of the same commit path, so a transition cannot be applied without being recorded. Recovery rebuilds the in-memory scheduler view from the relational store, then reconciles against the log to find transitions that were recorded but not applied.
| Option | Verdict | Reasoning |
|---|---|---|
| Relational authoritative for scheduling, event log for history, stated precedence | Chosen | Each consumer reads the store that fits it, and the tie-break is written down before it is needed |
| Event log as the single source of truth, state folded on read | Rejected | Elegant and immutable; folding for every dispatch decision at 3,200 jobs a minute is the wrong hot path |
| Relational only, history in an audit table | Rejected | Simpler; the audit table is mutable and shares the store's failure domain, so it fails when it is needed |
| The runner's own report as authoritative for its job | Rejected | Removes a store and makes a partition indistinguishable from a success — see ADR-13 |
What it buys
- Dispatch decisions are transactional and fast, and history is complete and immutable
- Recovery has a defined procedure rather than an argument, because precedence was decided in advance
- Losing the metadata store loses current state, not the account of what happened
What it costs
- Two stores to operate, replicate and reason about, with a reconciliation step in the recovery path
- Every transition costs a write to both, which is a throughput floor on the whole platform
- A bug that writes one and not the other produces a divergence that only the reconciliation finds
Choose differently when. At a fraction of this volume, folding the log on read would be entirely affordable and the second store unnecessary. Conversely, if scheduling state grew large enough that a single relational store became the bottleneck, the answer is sharding it rather than collapsing back to one store.
Why it holds up over time. The separation of current state from the account of how it got there is a distinction, not a technology, and it survives every substitution of database or stream. Naming the precedence rule is the part that pays off, because that is what an on-call engineer needs at 03:00 and cannot derive.
Lesson. Two stores are fine. Two stores without a written precedence rule are one incident away from being a guess.
ADR-13 · Runner loss is detected by lease expiry, and a job is never successful on an absent runner
Status: Accepted · Shown on views: 13, 19, 22
A job has stopped reporting. Is it dead, slow, or finished successfully with a lost response — and what does the platform do about it?
Context. This is the network-partition question in its most concrete form, and CI platforms get it wrong in a characteristic way: they treat the absence of a failure report as evidence of success, or they treat the absence of a heartbeat as evidence of death and re-dispatch a job that is in fact still running. Both are bad, but they are not equally bad. Re-dispatching a live job costs capacity and, for a job with side effects, may duplicate them. Marking an absent job successful puts an unverified green check on a pull request, which is the failure this platform exists to prevent. The asymmetry decides the design.
Decision. Every dispatched job holds a lease that it must renew. Loss is detected by lease expiry, never by silence on the result channel. An expired lease means the job is re-dispatched with the attempt count incremented and the cause recorded. A job is never reported successful on the basis of an absent runner. Automatic retry happens only for platform failure or a declared-retryable infrastructure error, never for user failure, and every retry records its cause.
How it is realised on AWS. The scheduler issues a lease with each job assignment and the per-job agent renews it; expiry re-queues the job. Partial logs already streamed are preserved and attached to the failed attempt, so the engineer sees what happened before the loss. Because the artefact is sealed by the attestor after the report, a job whose report is lost has no attestation and therefore cannot be treated as complete even if its bits were uploaded.
| Option | Verdict | Reasoning |
|---|---|---|
| Lease expiry detects loss; never successful on absence; retry only platform failure | Chosen | Resolves a partition to re-dispatch, which is the safe side of the asymmetry |
| Heartbeat timeout with optimistic completion on late report | Rejected | Reduces wasted work and occasionally produces a green check nobody earned |
| Trust the runner's report, no lease | Rejected | Simplest; a silent runner becomes indistinguishable from a passing build |
| Retry every failure a fixed number of times | Rejected | Hides real failures, wastes capacity on compile errors, and makes the infrastructure-failure ratio unmeasurable |
What it buys
- A partition costs duplicated work rather than a false green, which is the correct side to fail on
- Retry cause is recorded, which is what makes the infrastructure-failure ratio meaningful rather than decorative
- Partial logs survive a lost runner, so the engineer has something to read
What it costs
- A live-but-partitioned job is re-dispatched and its work duplicated, occasionally twice
- Jobs with external side effects need idempotency of their own; the platform cannot provide it for them
- Lease renewal is traffic on the hot path of every job, and its own availability matters
Choose differently when. If jobs were long, expensive and side-effecting — a multi-hour deploy rather than a three-minute test — the cost of duplicating one would exceed the cost of waiting longer, and a design with a longer lease plus explicit operator adjudication would be better. At a median job of 3m 10s, re-dispatch is the cheap option.
Why it holds up over time. "Absence of evidence is not evidence of success" is a distributed-systems invariant, and leases have been the mechanism for expressing it for decades. What changes is the timeout tuning; the asymmetry that decides the design does not.
Lesson. When you cannot distinguish two states, choose which error you would rather make and design for it explicitly. Here, wasting a build is always better than blessing one.
ADR-14 · Six job outcomes, and a platform failure is never reported as the user's
Status: Accepted · Shown on views: 04, 18, 22
What does a red check actually mean, and can the engineer tell from it whether the problem is theirs?
Context. Most CI systems report two outcomes: pass and fail. Everything that is not a passing test — an image pull timeout, a reclaimed instance, a registry outage, a cancelled run, a policy denial — arrives as fail, and the engineer's only recourse is to read the log and guess. The consequence is cultural rather than technical. An engineer who cannot distinguish a flake or an infrastructure blip from their own regression learns that the cheapest response to red is to press retry, and after that they retry genuine failures too. At that point the platform has trained its users out of reading its output, and no amount of dashboard investment recovers it.
Decision. A job reports exactly one of six outcomes: success, user failure, platform failure, cancelled, timed out, or blocked by policy. A platform failure is never attributed to the user's code. The infrastructure-failure ratio — platform failures as a share of all jobs — is a headline SLO at ≤ 0.5% weekly, alerting at 1.0% hourly. Flake rate, defined as differing outcomes on an identical commit and definition digest, is measured and reported per pipeline but deliberately not bounded: detection is the platform's obligation, fixing the test is the owning team's.
How it is realised on AWS. The per-job agent and the sandbox manager classify at the point of failure, where the cause is known, rather than post-hoc from a log. Retry policy keys off the class. The run console surfaces the class, the cause and a first-failure summary — the failing job, step and the log region around the first error. Flake detection compares outcomes across attempts on an identical commit and definition digest.
| Option | Verdict | Reasoning |
|---|---|---|
| Six outcomes, retry keyed off class, infra-failure ratio as a headline SLO | Chosen | Makes trust in a red build measurable, and makes retry a decision rather than a reflex |
| Pass/fail with the cause in the log | Rejected | Universal and cheap; it is how platforms teach engineers to stop reading |
| Pass/fail/error, three outcomes | Rejected | Better, and still conflates cancellation, timeout and policy denial with infrastructure error |
| Automatic flake quarantine after N differing outcomes | Deferred | Attractive and hazardous: quarantining a test is how a real regression ships. Report it, do not disable it |
What it buys
- An engineer can tell from the outcome alone whether the problem is theirs, which is what keeps a red build worth reading
- The infrastructure-failure ratio becomes measurable, and it is the number that decides whether the platform is trusted
- Retry stops being a cultural reflex and becomes a classified, recorded action
What it costs
- Every failure path in agent and manager must classify correctly, which is real work spread across the whole codebase
- Six outcomes are more than a badge can express, so the surfaces that show them need design attention
- A misclassification is worse than no classification, because it is confidently wrong
Choose differently when. If the platform served a single team that reads every log anyway, the taxonomy would be documentation rather than mechanism. It earns its cost at 4,200 engineers who will never read the platform's documentation and will absolutely learn a reflex.
Why it holds up over time. The distinction between "your change is wrong" and "our platform failed you" is about accountability, not technology, and it will matter identically on any future substrate. Designing the outcome taxonomy so the ratio is measurable is a permanent asset that survives every migration.
Lesson. The output of a build system is not a boolean; it is an attribution. Get the attribution wrong and users stop reading the output, which costs more than any latency regression.
Promotion and gating
The decisions that make "did we ship what we tested" answerable, and the gates unweakenable by the author.
ADR-15 · Build once, promote the digest, bind configuration at deployment
Status: Accepted · Shown on views: 06, 14, 17
Does the artefact that reaches production come from the run that passed staging, or from a fresh build of the same commit?
Context. Rebuilding per environment is intuitive and common: the same commit, built again with the environment's settings, gives an artefact tailored to its target. It also breaks the only chain of evidence linking what was tested to what is running, because the production artefact has never been tested — an artefact built from the same source is not the same artefact unless the build is bit-for-bit reproducible, and almost no real build is. Timestamps, dependency resolution drift, base-image tag movement and toolchain patch levels all vary between two builds minutes apart. The counter-pressure is that build-time configuration is genuinely convenient, and teams will have baked environment assumptions into their images for years.
Decision. One artefact digest is produced once and promoted, unchanged, through every environment. Environment-specific configuration is bound at deployment time, not at build time. Promotion moves a pointer; the bits never change. Rollback repoints to a retained previously-deployed digest and is an ordinary operation with a latency target rather than a recovery procedure.
How it is realised on AWS. Artefacts are content-addressed, and an environment holds a pointer to one digest plus its configuration and gates. The promotion path is ordered per pipeline, and skipping a stage requires a recorded break-glass override with actor, reason and expiry. A deployment lock serialises promotions to one environment. Retention keeps prior production digests so rollback has somewhere to point.
| Option | Verdict | Reasoning |
|---|---|---|
| Build once, promote by digest, configuration at deploy time | Chosen | "Did we ship what we tested" gets a cryptographic answer, and rollback becomes a pointer move |
| Rebuild per environment from the same commit | Rejected | Convenient for build-time configuration, and the production artefact has never been tested |
| Build once, then re-tag and re-sign per environment | Rejected | Keeps the bits and adds a mutable signing step that weakens the provenance for no gain |
| Reproducible builds so a rebuild is provably equivalent | Deferred | Would make rebuilding safe and is a multi-year effort across every language in the estate; a good ambition, not a plan |
What it buys
- The production artefact is byte-identical to the tested one, and the attestation covers exactly those bits
- Rollback is a five-minute pointer move rather than a twenty-minute rebuild under pressure
- Audit becomes a query: one digest, one attestation, one chain of gate evaluations
What it costs
- Build-time configuration must be migrated to deployment-time binding, which is visible work for every team
- Prior digests must be retained per environment so rollback has a target, which is storage the platform pays for
- One artefact must be able to run in every environment, which constrains how images may be built
Choose differently when. If builds were genuinely bit-for-bit reproducible across the whole estate, rebuilding per environment would be provably equivalent and the constraint on build-time configuration could be relaxed. That is the deferred option, and it is the only condition that changes this decision.
Why it holds up over time. Content addressing plus a mutable pointer is the pattern every artefact ecosystem has converged on, from container registries to package managers to source control itself. It outlives any particular registry because the identity comes from the bytes rather than from the store.
Lesson. If two things must be the same, do not build them twice and compare. Build once and move the result — and then the sameness needs no argument.
ADR-16 · Gates are evaluated outside the pipeline they govern, and fail closed in production
Status: Accepted · Shown on views: 06, 17, 18
Can the author of a pipeline weaken the checks that decide whether their own change reaches production?
Context. Gate jobs inside the pipeline are the natural design: the checks live beside the code they check, they are versioned with it, and a team can evolve them. They are also self-certification. The pipeline definition is authored by the same people whose change is being gated, so a gate expressed as a pipeline job can be reordered, made conditional, marked continue-on-error, or removed — usually with an entirely sincere reason, under deadline pressure, in a diff nobody reviews closely because it is 'just CI config'. The second half of the question is what happens when the external decision service is unavailable: failing open keeps delivery moving and silently removes every control, and failing closed stops delivery during an outage of something that is not the application.
Decision. Gates are evaluated by a decision service outside the pipeline, consulted by the deployer. A pipeline author cannot weaken the gates applying to their deployment. The service returns allow, deny, or allow-with-override-required, always with a machine-readable reason naming the failing condition. Production fails closed when the service is unavailable; a tenant may configure fail-open for non-production only. Overrides are recorded with actor, reason and expiry, and override usage is a watched signal.
How it is realised on AWS. Policy bundles are versioned per repository and held in the control plane, cached per tenant for a p99 decision under 500 ms. The deployer consults the service before moving an environment pointer; the verdict and its reason are written to the immutable audit trail before the pointer moves. Gate types include required job outcomes, code-owner and human approval, provenance and signature verification, vulnerability-severity thresholds, change-freeze windows, target SLO health and a required soak. Separation of duties is enforced where tenant policy demands it.
| Option | Verdict | Reasoning |
|---|---|---|
| External decision service, fail closed in production, recorded overrides | Chosen | The only design in which a gate is not under the control of the change it gates |
| Gate jobs inside the pipeline | Rejected | Simplest and most flexible, and is self-certification wearing an automation costume |
| Gates in the target platform's admission layer | Right elsewhere | Excellent for runtime invariants and blind to build-time evidence; complementary rather than alternative |
| External service with fail-open everywhere | Rejected | Keeps delivery moving during an outage by removing every control at exactly the moment nobody is watching |
What it buys
- A gate means the same thing on every pipeline, and changing it is a policy change with an owner
- Every verdict, approval and override has an audit record attributable to an identity
- Provenance verification becomes enforceable rather than advisory, because the enforcer is not the build
What it costs
- Every promotion depends on a service that must be more available than the deployments it governs
- Teams lose the ability to express a local, sensible gate as a pipeline job and must go through policy
- Break-glass override is necessary and is the control most likely to be abused
Choose differently when. If every team were both the author and the accountable owner of its production environment, with no separation-of-duties requirement and no shared blast radius, in-pipeline gates would be adequate and much simpler. The decision is driven by nine business units sharing a platform and by audit obligations that require the gate to be outside the gated.
Why it holds up over time. Separating the decision from the thing being decided is a governance invariant, and the specific gate types will churn constantly — vulnerability thresholds, freeze windows, SLO health, whatever comes next — without disturbing the boundary. That is exactly why the boundary rather than the gate list is what this record commits to.
Lesson. A control the controlled party can edit is not a control. Put the decision somewhere the change cannot reach, and make the override expensive to use rather than impossible.
Cost, retention and developer trust
The decisions that keep the safe path the cheap path, and a red build worth reading.
ADR-17 · Retention is set by artefact provenance class and pipeline class, not globally
Status: Accepted · Shown on views: 10, 11, 18
How long is a build artefact, a log and an attestation kept — and is the answer the same for a pull-request build as for a release?
Context. A global retention policy is easy to state and always wrong in one of two directions. Set it long enough for audit and the platform pays to keep 480 TB of pull-request artefacts and 6 TB a day of logs from builds nobody will ever look at again. Set it short enough to be affordable and the evidence an auditor needs is gone. The volumes make this the platform's fastest-growing cost line and the one most easily wasted, because nobody notices over-retention until the bill arrives, and by then the data cannot be un-kept cheaply.
Decision. Retention is a property of what produced the data. Release artefacts are kept 400 days, main-branch artefacts 90 days, pull-request artefacts 14 days. Logs are searchable 14 days, retrievable 90 days, and archived 400 days for release pipelines only. Caches are evicted after 7 days idle. Provenance attestations, the transparency log and the audit trail are kept 7 years and are immutable. Run metadata is kept 400 days. Eviction is auditable, and nothing referenced by a live environment pointer or a retained attestation is ever evicted.
How it is realised on AWS. Retention class is written on the artefact at seal time from the run's trigger and ref, so eviction never has to re-derive intent later. Object lifecycle policies implement the tiers; the immutable classes sit under an object-lock policy in the custody zone. Cost per job-minute and artefact TB by class are reported per tenant so the expensive habits are visible to the teams that have them.
| Option | Verdict | Reasoning |
|---|---|---|
| Retention by provenance and pipeline class, immutable evidence separate | Chosen | Spends storage where it is asked for and reclaims it where it never will be |
| One global retention period for everything | Rejected | Either unaffordable or non-compliant, and usually both at different layers |
| Per-tenant retention chosen by each team | Rejected | Every team chooses the maximum, because nobody pays for storage they do not see |
| Usage-driven retention — keep what has been accessed recently | Deferred | Efficient for artefacts and dangerous for evidence, which is by definition accessed rarely and needed absolutely |
What it buys
- The largest volumes are the shortest-lived, so the cost line tracks activity rather than accumulating forever
- Audit obligations are met by a small, immutable, clearly-bounded set of data
- Retention class is decided at seal time from facts, so eviction never needs to guess at intent
What it costs
- More classes to reason about, and a wrongly-classed artefact is either wasteful or gone too soon
- Pull-request artefacts vanish in 14 days, which occasionally frustrates a long-running investigation
- The 7-year immutable set cannot be trimmed, so a classification error there is permanent
Choose differently when. If artefact and log volumes were small enough that storage was not a material cost, a single long global retention would be simpler and harmless. At 480 TB growing 40% a year and 6 TB of logs a day, it is the fastest way to waste money on data nobody wants.
Why it holds up over time. Tying retention to the provenance of the data rather than to a global default is a policy shape, not a storage feature, and it transfers to any storage technology. As volumes grow, the argument for it strengthens; it never weakens.
Lesson. Retention is not one number. Decide it from what produced the data, write the class down at the moment you still know why, and make the cost visible to whoever generates it.
Every package used, in one table
The terms below are used precisely in this package. Several are used loosely in the wider CI/CD literature, and the difference matters when reading the decision records.
| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Run | One execution of a pipeline definition for one trigger event, holding a DAG of jobs. | The unit of admission, trust classification and audit. | A 'build', which in common usage means either a run or a single job and so is avoided here. |
| Job | One node of a run's DAG, executed in exactly one sandbox. | The unit of scheduling, isolation, credential scope and outcome reporting. | A 'step', which is a command inside a job and shares neither its sandbox boundary nor its outcome. |
| Trust class | Trusted or untrusted, computed at admission from the trigger event and membership records. | The single bit that decides credentials, cache write authority and publish identity. | A permission declared in the pipeline, which would let the code under test decide how far it is trusted. |
| Effective definition | The fully-resolved pipeline definition — templates expanded, digests pinned, parameters bound — recorded immutably per run. | What actually ran, as opposed to what the repository says today. | The pipeline file at HEAD, which is mutable and therefore useless as a run record. |
| Sandbox | The single-use, hardware-isolated environment created for one job and destroyed after it. | The trust boundary; the only place tenant code executes. | A 'runner', which in most CI vocabularies is a long-lived host that executes many jobs. |
| Attestation | A signed statement composed by the control-plane attestor naming the commit, definition digest, dependencies and builder for one artefact digest. | The non-forgeable link between source and deployed bits. | A build-signed statement, which is a claim by the thing being described. |
| Digest | The content hash of an artefact's bytes, and its permanent identity. | What is promoted, verified, deployed and rolled back to. | A tag or version, which is a mutable pointer to a digest. |
| Environment | A named, governed target holding a pointer to one digest, its configuration, its approvers and its gates. | The unit of promotion, locking and rollback. | A cluster or namespace, which is where an environment happens to run. |
| Gate | A condition evaluated outside the pipeline before an environment pointer may move. | The enforceable control on what reaches production. | A gate job inside the pipeline, which the pipeline's author can weaken. |
| Lease | A time-bounded claim a dispatched job must renew. | How runner loss is detected, and why a partition resolves to re-dispatch rather than to success. | A heartbeat, whose absence is often read as failure or, worse, as completion. |
| Platform failure | A job outcome caused by the platform rather than by the code under test. | The class that is retried automatically and counted in the infrastructure-failure ratio. | A generic 'error', which conflates infrastructure, cancellation, timeout and policy denial. |
| Infrastructure-failure ratio | Platform failures as a share of all jobs, measured weekly. | The single number that decides whether engineers read a red build or reflexively retry it. | Build success rate, which mixes the platform's reliability with the code's quality. |