Streaming Video Encoding & Packaging Pipeline

Architecture Views

22 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

Read it in seven acts. Acts 1 and 2 set the boundary and say what a viewer and a content operator actually get. Acts 3 to 5 are the structure, the data and what happens when a master arrives. Acts 6 and 7 are how it is operated and why it is safe. One decision runs through all of them: the master plus the recipe is the asset, and every rendition is a cache entry.

Context and scope

What sits inside the boundary — master in, published manifest out — and what is deliberately consumed rather than owned.

People and journeys

Who the pipeline is for, and the three journeys that decide whether it is any good: a premiere, a long-tail title, and a title that has to be on air by morning.
03 People who watch Viewer on a big screen assumed 40 M MAU Goal — I press play on the thing everyone is talking about and it looks like television, immediately, with no spinner and no soft first ten seconds. Core journeys Watch a new release premiere night Resume something old Viewer in the long tail 72% of catalogue hours Goal — I picked a 1974 documentary nobody else is watching tonight, and I would rather not be punished for having unusual taste. Core journeys Play a cold-stored title rendition not resident People who ship the content Content operations assumed 24 staff Goal — I need tomorrow's premiere playable on every device before the embargo lifts, and I need to know hours in advance if it will not be. Core journeys Get a title to air 1,800 versions/day Waive a narrow QC failure Video quality engineer owns the gate Goal — I want the gate to fail the things viewers would notice, and I want to know which encoder build caused a regression without bisecting the catalogue. Core journeys Tune a ladder from telemetry Roll out a new encoder build Media FinOps cost per delivered hour Goal — I want to see what a codec or a rung costs before it is switched on for 400,000 titles, not in next month's invoice. Core journeys Price a codec addition Partners, systems and machines Rights holder assumed 310 suppliers Goal — I hand over one master in the format we agreed and get told within the hour whether you accepted it, not three days later. Core journeys Deliver and redeliver a master DRM licence service 3 systems Goal — Issue a licence for a key the pipeline never stored itself. Core journeys Serve a playback licence Platform SRE on call Goal — When 60% of the fleet is taken away at 03:00 I want the queue to drain anyway and nobody to page me. Core journeys Survive a reclaim storm Actors and Their Core Journeys Goals are in the actor's voice. Journeys with an id get their own map; the rest are named so the set cannot pretend they do not exist. v 1.0 · owner Media Platform Architecture Actors and Their Core Journeys Eight actors, their goals in their own words, and which journeys get their own map. HTML page SVG draw.io

Structure

The parts, the planes they belong to, and the interfaces the platform exposes in both directions.

Data

What is irreplaceable, what is recoverable, what is deliberately disposable, and the lineage chain that survives deleting the bytes.

Runtime

What actually happens: one master to one published manifest, a chunk lost to a reclaim, and a rendition that does not exist when a viewer asks for it.

Operations

Rolling out an encoder build without re-encoding the catalogue, where everything runs, what is watched, and the lifecycle a rendition moves through.

Assurance

Why an unreleased master and a content key are safe, and what the design assumes will fail.
22 Assumed failure Detected by Immediate handling Residual exposure Capacity reclaim 60% of fleet in 10 min Lease expiry Requeue one chunk None — expected state Poison source Decoder hangs or crashes Chunk timeout + memory cap Quarantine with diagnosis Title misses its deadline Encoder build drift Worker ≠ recipe build Admission check Worker refuses task Throughput dip on rollout Stitch defect Visible pumping at a join Boundary verification Re-encode chunk pair Subtle artefacts below threshold Gate false pass Scores well, plays badly Playback telemetry Quarantine + manifest revision ≤0.2% of titles per quarter Key service outage KMS or DRM unavailable Key fetch latency alarm Stall publish, retry Premiere slips; never unencrypted Partial publish Some renditions absent Durability pre-check Publish narrower manifest Lower top rung on day one Cold-title demand spike Evicted rungs, sudden load Miss rate by title Coalesce + promote tier First viewers get a lower rung Region loss Encode region unavailable Control-plane health Fail over, regenerate Degraded rung set until refilled Cost runaway Retry loop or bad recipe Spend ceiling at admission Hard stop, not an alert Blocked job needs a human Failure Classes and Their Handling The residual column is the honest one: every row leaves something, and in eight of ten rows what it leaves is quality or lateness, never protection. v 1.0 · owner Media Platform Architecture Failure Classes Ten assumed failures, how each is detected, and what each one still leaves behind. HTML page SVG draw.io

Architecture One-Pager

A pipeline that turns one studio master into the few hundred files a player actually downloads — and decides, per title, how many of them are worth making.

The master plus the recipe is the asset. Every rendition is a cache entry. Everything else in this architecture is a consequence of that sentence.

Almost nobody outside the industry has a name for this system, and almost everybody has met it. A show lands at midnight and plays instantly in 4K on the television, while the same episode on a phone starts soft and sharpens eight seconds in. A film from 1974 looks worse than it should on a big screen; a cartoon from last week looks flawless at a quarter of the bitrate. A new release stalls for everyone in the first hour and is fine by morning. None of that is the network and none of it is the player. It is the encoding and packaging pipeline, and the reason it is hard is not video compression. It is that the catalogue is enormous and the viewing is not: an assumed 400,000 titles and 250,000 source hours, of which an assumed 72% of catalogue hours attract under 0.1% of total watch time. Pre-encoding a full ladder of a dozen renditions in several codecs for every one of them is a storage bill paid, in perpetuity, for files nobody requests. Not pre-encoding them means a real viewer waits. The architecture is the answer to where that line goes.

Seven stages, two execution fleets and three storage zones. A rights holder delivers a mezzanine master; the platform verifies it against a delivery manifest, probes it, classifies it against a conformance policy, and lands it write-once in dual-region archive storage. It then analyses the content — shots, spatial and temporal complexity, grain, motion, effective resolution after any upscale in the master — and solves a bitrate ladder for that title, recording the ladder, every encoder parameter and the encoder build identifier as a single immutable, content-addressed recipe. The recipe, not the rendition, is the authoritative derived artefact. A control plane on Spanner materialises the recipe into chunk tasks on shot boundaries and dispatches them to a Spot fleet whose workers are expected to be withdrawn continuously; a reclaim costs one chunk and is rescheduled without an operator. Stitched renditions pass a blocking gate — a full-reference perceptual score per rung, structural checks, and decodability on a device matrix — before a packager segments and encrypts them once under common encryption, with content keys held in memory for the duration of one packaging job and never persisted. Publication is an atomic swap of a manifest pointer, launch subset first. After publication, a lifecycle controller demotes and evicts renditions on an economic rule, and a separate warm fleet regenerates an evicted rendition when a viewer asks for it, falling back a rung rather than making anyone wait past the budget.

What it is, and what it is not

A pipeline that makes a master playable on every supported deviceA video compression research project
A system whose authoritative artefacts are a master and a recipeA rendition library that must be preserved
An economic argument about which renditions are worth existingA quality-at-any-cost encoder farm
A platform that proves quality with a number before publishingA workflow that relies on someone watching the output
Designed around compute being taken away continuouslyA highly available encode fleet
A publisher of immutable segments and mutable manifestsA CDN or an edge cache
A consumer of a DRM licence serviceA DRM or licence-issuing system
Scoped to video-on-demandA live or contribution encoding path
A multiplexer of audio and subtitle tracks delivered to itA subtitle authoring or translation tool

The decisions that are the architecture

01The master and the recipe are the only assets

Renditions get no backup, no cross-region replication and no RPO. Storage cost becomes a retention-versus-regeneration calculation per title, and disaster recovery becomes replication of the two small things plus a regeneration budget.

ADR-01

02The mezzanine is immutable on arrival

Write-once, retention-locked, never modified in place. A redelivery is a new source version requiring explicit promotion, not an overwrite — so a supplier cannot silently change what was published.

ADR-02

03The recipe is the only sanctioned encode input

Immutable, content-addressed, pinned to an encoder build. It is the join key between a master and everything derived from it, and the reason an encode can be re-run identically in eighteen months.

ADR-03

04Renditions are identified by content hash

A digest over source version, recipe and encoder build. An encoder upgrade therefore becomes a catalogue-wide cache miss filled lazily by demand, rather than a scheduled re-encode of 250,000 hours.

ADR-04

05Bound reclaim loss rather than checkpoint

Shot-boundary chunks sized so a withdrawn worker costs at most one chunk-minute, with chunk duration as tunable policy. No encoder state is serialised or revived.

ADR-05

06Determinism is a release gate

An encoder build that cannot replay a golden corpus byte-identically is rejected. Determinism is what makes retries free, caches trustworthy and lineage true rather than aspirational.

ADR-06

07Lanes with capacity floors, ordered by deadline

A premiere cannot be starved by a backfill, and a backfill cannot be starved indefinitely. Priority is derived from the deadline and the work remaining, not from a static field that is wrong by the second day.

ADR-07

08Batch and on-demand share code, never capacity

One is throughput-bound on reclaimable compute; the other is latency-bound with reserved headroom. Mixing them would either give a waiting viewer a Spot queue or give the batch fleet an availability requirement it does not need.

ADR-08

09The quality gate blocks, and waivers are on the record

A rendition below the floor is not publishable. The only route past it is a recorded waiver by an authorised approver, and the person under deadline pressure is not the person who sets the floor.

ADR-09

10Playback telemetry is a second gate

A rendition that scores well and plays badly on a real device is quarantined after publication and the manifest revised. Pre-publication metrics are necessary and not sufficient.

ADR-10

11Atomic manifest swap is the only publication primitive

An incomplete ladder publishes as a narrower manifest or not at all. Rollback is a pointer move completing in an assumed 120 seconds with no dependency on the encode fleet.

ADR-11

12The ladder is solved per title

Derived from a complexity profile, with the egress saved weighed against the analysis compute spent and the comparison recorded. At an assumed 40 PB/day of delivered egress, one percent of average bitrate is worth about 400 TB/day.

ADR-12

13Residency is an economic decision, re-made continuously

A rendition is demoted and eventually evicted when its regeneration cost falls below its retention cost over the policy horizon. Eviction is not an end state: a request re-enters the lifecycle at generation.

ADR-13

14Recovery is regeneration, not replication

The standby region carries a warm control plane and no encode fleet. After a region loss the published manifest degrades to the surviving rung set and refills — a visible consequence rather than a hidden one.

ADR-14

15The encode fleet is an untrusted zone

A container escape through a codec parser handling a hostile master is the assumed attack. Workers decode sandboxed with no egress, read one prefix and write one, and content keys never persist in the pipeline.

ADR-15

Why this should still hold up in ten years

Codecs, encoders and cloud price lists will all change inside the life of this design. The decisions above were chosen to be the ones that do not.

The long tail is a structural property of catalogues, not a fact about 2026

Viewing concentrates on a small fraction of any large catalogue, and catalogues grow faster than viewing hours per subscriber. The argument for treating renditions as a cache gets stronger as the catalogue grows, not weaker.

A new codec is a recipe policy change, not a migration

Because identity is a hash over source, recipe and encoder build, adopting a codec means new recipes and lazy regeneration. The design that pre-encodes everything has to schedule and fund a catalogue sweep for each new codec; this one does not.

Reclaimable compute is getting cheaper relative to reserved compute, not dearer

Every provider prices interruptible capacity below committed capacity because it monetises idle fleet. A design whose resilience comes from task granularity rather than worker protection keeps collecting that discount as the gap widens.

Determinism and content addressing are what let the system be re-reasoned about later

The expensive failure mode in a ten-year-old media pipeline is not a bad encode; it is a published byte nobody can explain. Recording the recipe and the build digest is cheap now and is the only thing that makes an audit, a regression hunt or a regeneration possible then.

The one invariant that must not be traded is protection

Quality, latency and rung availability are all allowed to degrade under failure, and every failure class in this design degrades one of them. Nothing degrades encryption, because that is the trade whose cost arrives as a contract breach rather than a support ticket.

Non-functional targets

Every target below is a stated assumption, invented for a large consumer video-on-demand service so that the architecture has something to be wrong about. The view column points at the diagram where the mechanism that delivers it is drawn.

QualityTargetHow it is metView
Control-plane availability ≥ 99.95% / month Regional Cloud Run and GKE across three zones; Spanner regional; no dependency on the encode fleet 17
Publish path availability ≥ 99.9% / month Manifest swap is a pointer move in the catalogue projection, independent of packaging and encode 13
Rollback time ≤ 120 s Previous manifest revision retained; rollback is a pointer move with no fleet dependency 19
Encode fleet availability none, by design Spot MIGs; up to 60% withdrawal in 10 minutes absorbed by chunk granularity 14
Catalogue-lane turnaround p95 ≤ 0.80× duration Chunked parallel encode across ~14,000 slots with per-lane capacity floors 14
Premiere launch subset ≤ 45 min for 120 min Three rungs, baseline codec, premiere lane floor; ≥ 160× real-time aggregate 04
Premiere full ladder p95 ≤ 4 h from ingest Remaining rungs added by later manifest revisions 19
On-demand first byte p95 ≤ 1.5 s, p99 ≤ 3.5 s Warm GKE pool, request coalescing, archive range read; nearest-rung fallback past budget 15
Master ingest volume ~126 TB/day 3,500 source hours/day at an assumed 80 Mbps mezzanine; write-once dual-region archive 10
Chunk task throughput 220 k/day, 600 k peak Pub/Sub dispatch with Spanner task claim; no shared mutable state on the encode path 14
Delivered egress ~40 PB/day Per-title ladder; 1% average bitrate ≈ 400 TB/day of egress 12
Top-rung quality ≥ 93 / 100 aggregate Full-reference perceptual score per rung against the source, blocking 18
Worst-window quality no 2 s window < 60 Worst-window score stored with the rendition, not just the aggregate 12
Gate false-pass rate ≤ 0.2% / quarter Post-publication telemetry as a second gate with automatic quarantine 22
Master and recipe durability RPO 0, 7-year lock Dual-region archive buckets with object retention locks; recipes dual-region 11
Control state recovery RPO ≤ 5 s, RTO ≤ 15 min Spanner with a warm standby control plane in the second region 17
Rendition recovery no RPO; refill 30 min p95 Deliberate: loss is a regeneration from master plus recipe 11
Full-ladder encode cost ≤ $1.10 / source hour Open-source encoders on per-second-billed Spot capacity 14
On-demand share of spend ≤ 6% of encode spend Budget check at admission; exceeding it re-tunes eviction rather than raising the cap 15

Scope

In scope

  • Delivery intake by resumable upload, supplier pull or object-store handoff, with integrity verification against a delivery manifest before any work is scheduled.
  • Deep technical probing and conformance classification — accepted, accepted-with-remediation, or rejected — with the applied remediations recorded against the source version.
  • Content complexity analysis and per-title bitrate ladder solving, including codec coverage as a per-title policy and refusal to upscale above the effective source resolution.
  • The recipe plane: immutable, content-addressed, encoder-build-pinned recipes as the only sanctioned encode input, with human overrides recorded with actor and reason.
  • Chunked parallel encode on reclaimable capacity with deterministic, idempotent chunk tasks, lane scheduling with capacity floors, quotas and per-title spend ceilings.
  • Stitching with boundary verification, and segment alignment across every rung of a title so a player may switch rungs at any segment.
  • A blocking quality gate: full-reference perceptual scoring per rung, structural checks, decodability on a device matrix, human sampling, and recorded waivers.
  • Packaging to segmented adaptive formats from one set of elementary streams, common encryption with keys from an external key service, and per-device-class manifests.
  • Atomic manifest publication, launch subsets, rollback, takedown with key revocation, and the demand-tiered residency, eviction and regeneration of every rendition.
  • Lineage from master version through recipe digest and encoder build to published byte, plus the append-only audit trail of overrides, waivers, publishes and rollbacks.

Explicitly out of scope

  • The CDN and edge cache — an adjacent use case in this practice.
  • The player, its adaptive bitrate logic and its device-side decoding.
  • Subtitle and audio-description authoring, translation and timing.
  • Catalogue, artwork and recommendation metadata systems.
  • Live capture, contribution encoding and near-live fast-turnaround paths.
  • The rights and licensing system that decides what may be published where.

What a prototype would have to prove

The decisions in this record are falsifiable, and most of them are falsifiable cheaply. A prototype on a few hundred source hours and a few hundred Spot cores settles the questions that would otherwise be discovered after thirty diagrams have been drawn and a year has been spent. These are the six things it should measure, and the seven scenarios it should survive.

  1. Determinism in practice: that the chosen encoder, wrapped and pinned, produces byte-identical output for the same source, recipe and build across machine types, kernel versions and restarts. If it does not, ADR-04, ADR-05 and ADR-06 all change.
  2. Boundary invisibility: the chunk duration at which stitched output stops showing quality pumping at joins, measured perceptually rather than asserted — this sets the floor for the reclaim-loss bound in ADR-05.
  3. Ladder value: the delivered bitrate saved by per-title solving against the analysis compute spent, on a corpus spanning grainy film, animation and sport, so the trade in ADR-12 has a measured exchange rate.
  4. Regeneration latency: the real archive-restore-plus-encode time for one segment of a top rung, which is the entire basis of the on-demand budget in ADR-08 and ADR-13.
  5. Reclaim behaviour: the actual withdrawal distribution for the chosen Spot machine family in the chosen region, because the chunk-size calculation is only as good as the interval it assumes.
  6. Gate agreement: how often the perceptual floor and a panel of human reviewers disagree, in both directions, which is what tells you whether the floor in ADR-09 is set anywhere near the right place.
  • Sixty percent of the fleet is withdrawn in ten minutes during a premiere encode. The job's progress must stay monotonic, no chunk may be lost twice, and the lane's p95 must hold.
  • A supplier redelivers a master for an already-published title. The published manifest must not change until someone explicitly promotes the new source version.
  • A new encoder build is promoted. No existing rendition may be re-encoded, no existing recipe may change, and the first new recipe must pin the new build.
  • A master is crafted to hang the decoder. The chunk must time out, retries must be bounded, the title must quarantine with a diagnosis, and the lane must keep draining.
  • The key service becomes unavailable mid-packaging. Publication must stall and retry, already-published content must be unaffected, and nothing unencrypted may reach any public path.
  • A long-tail title with every rung evicted becomes popular in one hour. Concurrent requests for the same missing rendition must produce one encode, and nobody may wait past the budget.
  • The encode region is lost. The control plane must fail over inside the RTO, in-flight jobs must restart from completed chunks, and the published manifest must degrade to the surviving rung set rather than break.

Open risks, carried rather than hidden

RiskIf it landsResponse
The demand prediction for a launch subset is wrong on day zero Premiere viewers get a lower rung in the first minutes — exactly the audience least willing to forgive it, and the one most likely to say so publicly. Keep the launch subset deliberately wider than the prediction suggests for titles above a marketing threshold, and treat the subset's width as a tunable with a cost attached rather than a model output. Question 3 stays open.
The chosen encoder turns out not to be deterministic across machine types Idempotent retries, content addressing and lineage all weaken at once; a duplicate execution stops being free and the cache stops being trustworthy. Make determinism a build-promotion gate before anything depends on it (ADR-06), and pin machine family as well as image digest if the gate cannot otherwise be met.
On-demand generation exceeds its spend ceiling routinely Either the cap is hit and viewers see fallbacks often, or the cap is raised and the economic case for eviction quietly disappears. Treat a sustained breach as a signal that the eviction policy is mis-tuned, not that the cap is wrong; the JIT share of spend is a first-class operational metric for that reason (view 18).
The perceptual gate passes renditions that real devices play badly Quality failures reach viewers, and the first people to notice are the audience rather than the platform. Post-publication telemetry as a second gate with automatic quarantine (ADR-10), and human sampling weighted towards narrow passes and new encoder builds.
A rights holder's contract forbids treating renditions as disposable The critical design decision does not apply to part of the catalogue, and a second, pre-encoded regime appears alongside the first. Make residency a per-title policy from the start so a contractual floor is a policy value rather than an architectural exception, and price the exception explicitly.
Archive-class restore latency makes the on-demand budget unreachable The cache model's cost advantage survives but its viewer-facing promise does not, and the fallback becomes the normal path rather than the exception. Measure restore latency in the prototype before committing; keep a hot copy of the master for titles above a demand threshold, which is a cheaper exception than keeping every rendition.

Architecture Decision Record

Why every component and every technology on these 22 views is what it is, and what each choice costs.

Fifteen decisions, grouped into five areas. Each carries the question that forced it, how it is realised on Google Cloud, what was rejected, what would flip it, and why it should still hold in ten years.

Status of this document. <b>Evidence note.</b> Every quantity in this package is a stated assumption for a large consumer video-on-demand service, invented so the architecture commits to something arguable. None is taken from a published figure for any named streaming service, and no claim is made that any of them reflects a real operator's numbers. Where a number drives a decision, the decision says so and names the number. Replace each one with a measured value before building anything from this.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

custody-and-identity 4

What the system actually owns, and how a derived artefact is named.

ADR-01The master plus the recipe is the asset; every rendition is a cache entry ADR-02The mezzanine is immutable on arrival, and a redelivery is a new source version ADR-03The recipe is the only sanctioned encode input, pinned to an encoder build ADR-04Renditions are identified by content hash, not by assigned identity

encode-execution 4

Turning a recipe into bytes on compute that is being taken away.

ADR-05Bound reclaim loss with small chunks rather than checkpointing encoder state ADR-06Determinism is an encoder release gate, not a hoped-for property ADR-07Priority lanes with capacity floors, ordered by deadline and work remaining ADR-08Batch and on-demand generation share code, never capacity

quality-and-publication 3

Proving an output is good enough, and putting it in front of people.

ADR-09The quality gate blocks publication, and a waiver is a recorded act ADR-10Playback telemetry is a second, post-publication gate ADR-11Atomic manifest swap is the only publication primitive

economics-and-lifecycle 2

Deciding which renditions are worth existing, and for how long.

ADR-12The bitrate ladder is solved per title from a complexity profile ADR-13Rendition residency is an economic decision, re-made continuously

protection-and-operations 2

Keeping an unreleased master and a content key safe, and recovering a region.

ADR-14Regional recovery is regeneration from custody, not rendition replication ADR-15The encode fleet is an untrusted zone, and no fallback weakens protection

Technology by capability

Google Cloud carries 7 of the 56 existing documents in this practice against Amazon Web Services' 10, open-source-on-premises' 11 and Microsoft Azure's 8, so it is the honest rotation — and the topic suits it. The central question here is cost per delivered hour, and that only becomes a real design question when the encoder is yours and the compute is reclaimable: per-second-billed preemptible instances, a managed batch scheduler, object storage with lifecycle classes, and a media-aware CDN downstream. A managed per-minute transcoding API would answer the interesting question with a price list. The requirement document stays vendor-neutral throughout; the choices below belong to this architecture, not to the ask, and each row names what it was chosen over.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Master custody — immutable, retention-locked, dual-region Cloud Storage dual-region bucket, Archive class, object retention lock Google Cloud Single-region with manual replication; a self-managed object store on-premises Dual-region gives RPO 0 without a replication job to operate, and retention lock is a storage-level guarantee rather than application logic the pipeline could be talked out of. ADR-02
Recipe store — immutable, content-addressed, indexed Cloud Storage dual-region for the documents, Spanner for the index Google Cloud A relational store holding the documents inline; a document database The documents are immutable and want object-store durability; only the lookup wants a transactional index. Splitting them keeps the custody artefact out of a schema that will change. ADR-03
Job and chunk-task state — transactional, cross-zone Cloud Spanner, regional Google Cloud PostgreSQL with a leader per region; a lease protocol over a key-value store This is the one place in the design needing transactional correctness across zones — task claim, lease expiry and lane accounting in one transaction. The alternative is a lease protocol we would have to write and prove. ADR-07
Batch encode fleet — reclaimable, per-second billed Compute Engine Spot VMs in zonal managed instance groups, gVisor around the decoder Google Cloud Committed-use reserved instances; a managed transcoding API priced per minute Reclaimable capacity is the economic premise of the whole design, and per-second billing is what makes small chunks affordable. A managed per-minute API removes the decision this use case exists to make. ADR-05
Encoders and packaging Open-source encoders (FFmpeg-family, SVT-AV1) and Shaka Packager on GKE Open source A managed transcoding and packaging service; a commercial encoder licence Owning the encoder is what makes determinism gateable, build pinning enforceable and the ladder solvable per title. Common encryption from one elementary-stream set is a packager requirement, not a service feature. ADR-06
On-demand generation — latency-bound, reserved GKE regional warm node pool beside a Cloud Run packaging origin Google Cloud Scale-to-zero serverless; preemption inside the batch fleet; edge compute An assumed p95 of 1.5 s to first byte rules out a cold start and rules out queuing behind Spot work. Edge compute is left open in Question 7 rather than chosen. ADR-08
Task dispatch — at-least-once, high fan-out Pub/Sub, with task claim in Spanner Google Cloud A broker with exactly-once semantics; direct scheduler-to-worker assignment Determinism and idempotent chunks make at-least-once sufficient, so the expensive delivery guarantee buys nothing. The transactional claim lives where the state already is. ADR-06
Content keys and encryption at rest Cloud KMS for CMEK and content-key generation, keys held in memory for one packaging job Google Cloud A self-hosted key manager; keys cached in the packaging tier for throughput The pipeline must never be a place a content key is stored, and per-rights-holder key separation is a contract requirement for part of the catalogue. Caching keys would trade a contractual exposure for a latency gain nobody asked for. ADR-15
Quality, cost and playback analytics BigQuery, partitioned by publish date, fed from Pub/Sub Google Cloud A time-series database; a warehouse in the analytics estate Gate decisions, metric scores and playback outcomes are append-only, queried by date and joined against digests — a columnar store is the natural shape, and the post-hoc gate needs to query it directly. ADR-10
Rendition residency and eviction Cloud Storage regional with lifecycle rules, driven by a demand-tier controller on GKE This design Pure age-based lifecycle rules; manual tiering; no eviction at all The eviction rule is economic rather than temporal — regeneration cost against retention cost over the policy horizon — and an age rule cannot express it. Lifecycle rules execute the decision; the controller makes it. ADR-13

The decisions, and the alternatives that lost

custody-and-identityWhat the system actually owns, and how a derived artefact is named.

ADR-01

The master plus the recipe is the asset; every rendition is a cache entry

Accepted

Is a published rendition a durable artefact the platform is obliged to preserve, or a derived file it may delete and rebuild whenever that is cheaper?

Context
A large catalogue produces a brutal asymmetry. An assumed 400,000 titles and 250,000 source hours, each wanting a dozen renditions across more than one codec, is on the order of 1.2 million rendition sets. An assumed 72% of catalogue hours attract under 0.1% of total watch time. Treating every rendition as a publication means paying regional storage, lifecycle management and — if they are treated as irreplaceable — cross-region replication and backup, in perpetuity, for files that will be requested a handful of times or never. The alternative is uncomfortable in a different way: if renditions can be deleted, a viewer can ask for one that does not exist, and somebody has to answer them inside a latency budget. Nothing else in the design can be settled before this is, because it determines what gets replicated, what gets an RPO, what disaster recovery means, what a codec upgrade costs, and whether there is an on-demand encoding fleet at all.
Decision
The mezzanine master and the recipe are the only assets. Renditions and segments carry no RPO, no backup and no cross-region replication; they are evictable under an economic policy and regenerable from the master and the recipe. On-demand regeneration is a first-class, latency-budgeted, separately scaled component of the architecture rather than a repair tool.
How it is realised on AWS
Masters land write-once in a dual-region archive bucket with object retention locks, and recipes sit alongside them in a dual-region bucket indexed in Spanner. Renditions and segments live in a single-region bucket, content-addressed, with lifecycle rules that demote and delete by demand tier. The storage-zones view draws exactly three zones — irreplaceable, recoverable, disposable — and the disposable zone's RPO is written on the diagram as "none, by design" so that nobody later adds replication to it believing they are fixing an oversight. A warm GKE pool beside the packaging origin serves regeneration with request coalescing on the rendition digest and a nearest-resident-rung fallback.
Options weighed
  • ChosenRenditions are a cache over master plus recipe: Storage becomes a per-title economic decision; DR becomes replication of two small things plus a regeneration budget; the cost is a viewer-visible fallback on cold titles.
  • Right elsewherePre-encode the full ladder and preserve every rendition: Right for a small catalogue, a premium archive, or an operator whose contracts forbid disposing of deliverables. Uniform latency, no fallback to explain, and a storage bill that scales with catalogue rather than viewing.
  • RejectedPre-encode everything but back up nothing: Keeps the full storage bill and still has to regenerate after a loss — the worst of both, with no on-demand path built because nobody planned for one.
  • RejectedEncode only on first request, with nothing pre-encoded: Makes the premiere the failure case. The first hour of a new release is precisely when nothing may be generated on demand.
Consequences
What it buys
  • Disaster recovery for 1.2 million rendition sets reduces to replicating masters and recipes — orders of magnitude smaller — plus a measured regeneration budget.
  • A bad encode is a recipe revision rather than a data-repair exercise, and a codec upgrade is a policy change applied lazily as demand arrives.
  • The long tail stops being a perpetual storage obligation and becomes a retention-versus-regeneration calculation that can be re-made as prices move.
  • Storage class, eviction threshold and launch-subset width all become tunable policy values with a visible cost, rather than architectural facts.
What it costs
  • A real viewer can experience a cache miss, and the honest p99 of that miss is an assumed 3.5 seconds to first byte.
  • A nearest-rung fallback is visible: somebody will occasionally watch a lower rung than their connection deserves, and that has to be acceptable to the product.
  • The architecture now needs a second encoding fleet with its own latency SLO, its own reserved headroom and its own spend ceiling.
  • Any rights holder contract that forbids disposing of deliverables becomes an exception that must be priced and policed per title.
Choose differently when
Three things would flip it. If storage became cheap enough relative to compute that retaining a full ladder for the whole catalogue cost less than the on-demand fleet plus its fallbacks, pre-encoding wins outright. If archive-class restore latency proves unreachable within the first-byte budget even for a warm pool, the cache model keeps its cost advantage but loses its viewer-facing promise, and the honest response is a hot master copy for high-demand titles rather than retaining renditions. And if a majority of the catalogue arrived under contracts that forbid treating deliverables as disposable, the exception would be the rule and the decision should be reversed rather than special-cased.
Why it holds up over time
The decision rests on a property of catalogues rather than a property of 2026: viewing concentrates, and catalogues grow faster than viewing hours per subscriber. As a catalogue grows the case strengthens. It also gets stronger as codecs multiply, because every new codec multiplies the rendition count under the pre-encode model while costing only new recipes under this one. The main thing that would date it is a step change in storage economics, and storage has historically fallen in price more slowly than codec counts have risen.
LessonDecide what is genuinely irreplaceable before deciding anything about durability. Most of what a pipeline produces is a function of something smaller, and naming that function is cheaper than preserving its output.
Shown on views11 15 19
ADR-02

The mezzanine is immutable on arrival, and a redelivery is a new source version

Accepted

When a supplier sends a corrected master for a title that is already published, does it replace the old one, or become a new thing that someone has to promote?

Context
Redelivery is routine: a wrong audio mix, a missing frame, a mastering error, a different territorial cut. The tempting behaviour is to overwrite, because the supplier thinks of it as a correction and the file has the same name. Overwriting breaks three things at once. It invalidates every rendition derived from the old bytes without any record that it did so; it makes the lineage claim false, because the published byte can no longer be traced to the master that produced it; and it means a supplier can change what viewers see without anyone at the platform deciding that they should.
Decision
Received bytes are immutable on arrival: never modified in place, never overwritten by a later delivery. Integrity is verified against a delivery manifest before any work is scheduled. A redelivery becomes a new source version, and replacing what is published requires an explicit promotion.
How it is realised on AWS
The upload endpoint writes to a write-once prefix in the dual-region archive bucket under a key derived from the content checksum, with object retention locks preventing modification for the retention term. The source_version row carries the checksum as a unique key, the full technical profile from the deep probe, the conformance verdict and the applied remediations. Promotion is a separate, audited operation that creates new recipes; it does not mutate the existing ones. Conformance runs before scheduling so a rejection reaches the supplier within an assumed one hour rather than after a day of encoding.
Options weighed
  • ChosenImmutable landing, redelivery as a new version, explicit promotion: Lineage survives, the platform decides what is published, and a bad redelivery cannot silently invalidate a working title.
  • RejectedOverwrite in place and re-encode: Cheap to implement and destroys traceability. Also makes rollback impossible, because the bytes the current renditions came from no longer exist.
  • RejectedVersioned bucket with the latest version implicitly authoritative: Keeps the history but still lets a supplier's upload change what is published. The storage layer should not hold an editorial decision.
  • RejectedAccept any delivery and let the quality gate catch problems: Spends a full ladder of encode on a master that a checksum would have rejected in seconds, and tells the supplier far too late to redeliver before air.
Consequences
What it buys
  • Every published byte traces to a specific master version, which is what makes a regression hunt or an audit possible years later.
  • Rollback is real: the previous source version and its recipes still exist, so reverting is a promotion rather than a recovery.
  • A redelivery cannot invalidate a live title as a side effect; somebody has to choose, and the choice is audited.
  • Rejecting on integrity before scheduling saves an assumed full-ladder encode on every malformed delivery and gets the supplier a verdict inside an hour.
What it costs
  • Storage holds every version of every master for the retention term, including the ones that were superseded the same week.
  • Operations carries a promotion step that someone has to perform, and a queue of unpromoted redeliveries that someone has to watch.
  • Suppliers who expect overwrite semantics need an explanation, and some will need a contract amendment.
  • The checksum-derived key means a byte-identical redelivery is a no-op, which is correct but occasionally confusing to a supplier who expects to see a new record.
Choose differently when
If retention cost for superseded masters became material — a catalogue dominated by high-bitrate redeliveries, say — a tiering rule that deletes unpromoted source versions after a window would be a reasonable amendment, keeping immutability while bounding the history. The decision would only genuinely flip if the platform stopped being the publisher of record, for instance if a supplier operated the publication decision themselves, at which point immutability on arrival is still right but promotion moves to them.
Why it holds up over time
Immutability on write has become the default expectation for object storage and for anything auditable, and retention locks are now a standard storage feature rather than an add-on. The decision gets easier to implement over time, not harder. The underlying reason — that an editorial decision should not be expressible as a file overwrite — does not depend on any technology.
LessonMake the system of record immutable, and make the act of changing what people see a separate, named, audited operation. Those two things are not the same and should not share a mechanism.
Shown on views10 12 06
ADR-03

The recipe is the only sanctioned encode input, pinned to an encoder build

Accepted

Where do the encoding parameters for a title live — in the job, in the worker's configuration, or in a separate immutable artefact the job merely references?

Context
Parameters have a way of accumulating in whichever place is most convenient to change: a job submission field, a per-worker config file, an environment variable set during an incident. Each of those makes the parameters un-reconstructable afterwards. Meanwhile the one derived artefact in the whole pipeline that cannot be recomputed from anything else is the decision about how a title should be encoded: the analysis that informed it, the ladder that came out, the codec policy, and the encoder build it was meant for. Everything else — chunks, renditions, segments, manifests — is a function of the master and that decision.
Decision
The recipe is an immutable, versioned, content-addressed document holding the ladder, the analysis inputs, every encoder parameter and the encoder build identifier. It is the only sanctioned input to an encode. No out-of-band parameter may reach a worker, and a human override is recorded with actor and reason or refused.
How it is realised on AWS
Recipes are JSON documents keyed by a digest over source version, analysis output, encoder build and parameters, stored in a dual-region bucket with a Spanner index. Chunk tasks carry the digest, not the parameters. A worker fetches the recipe by digest and refuses the task if its own encoder build does not match the pinned one. The layered-architecture view puts the recipe plane above the control plane and shades it, because its lifetime and its custody requirements are those of the master rather than those of a job.
Options weighed
  • ChosenImmutable content-addressed recipe, build-pinned, as the only input: Makes an encode reproducible years later and makes lineage a lookup rather than an archaeology exercise.
  • RejectedParameters carried on the job record: Job records get pruned, and parameters that live inside a mutable row stop being able to explain a published byte once the row changes.
  • RejectedParameters in worker configuration, deployed with the fleet: Makes an encode a function of when it ran. Two chunks of the same rendition can then differ, which is the defect that is hardest to find.
  • DeferredA ladder template per content category, resolved at encode time: A reasonable optimisation for the solver's inputs, but the resolved result must still be frozen into a recipe before any chunk runs.
Consequences
What it buys
  • An encode is reproducible from the recipe digest alone, which is what makes retries free and regeneration possible eighteen months later.
  • Lineage from published byte to master version is a two-hop lookup rather than an investigation.
  • An incident cannot be 'fixed' by changing a parameter on a worker, because the worker has no parameters to change.
  • Overrides become visible: the audit log shows who widened a ladder and why, which is the only way to learn whether overrides are a symptom.
What it costs
  • Every parameter change produces a new recipe version, so the recipe store grows with experimentation as well as with the catalogue.
  • The solver has to be complete: anything it forgets to record is a parameter that cannot be reconstructed, and the gap will not be obvious.
  • Operators lose the ability to make a quick fleet-wide tweak, which is occasionally genuinely inconvenient during an incident.
  • A build-pinned recipe means the fleet must be able to run older builds, so build images have to be retained as long as the recipes that reference them.
Choose differently when
If encoder builds could not be retained for the life of their recipes — a licensing restriction, or an upstream project that deletes releases — build pinning becomes unenforceable and the honest fallback is to record the build identifier descriptively and accept that regeneration may be equivalent rather than identical. That weakens ADR-04 and ADR-06 with it, so it should be resisted rather than accommodated quietly.
Why it holds up over time
Content-addressed immutable configuration is the direction every build, deployment and data system has moved for two decades, for the same reason: it is the only way to answer what produced a given output. The specific parameters will change with every codec generation; the decision to freeze them in an addressable artefact will not.
LessonFind the one derived artefact in your pipeline that cannot be recomputed, give it custody and an address, and make everything else a function of it.
Shown on views07 12 16
ADR-04

Renditions are identified by content hash, not by assigned identity

Accepted

Is a rendition named by a digest of what produced it, or by an identifier the platform assigns to the slot it fills?

Context
This looks like a naming convention and is actually the decision that prices every future codec and encoder migration. Under a content hash over source version, recipe and encoder build, promoting a new encoder build changes the digest of everything it would produce, which means the whole catalogue becomes a cache miss. Under an assigned identity — title, rung, codec — the upgrade is invisible: the same slot is simply filled by newer bytes next time, and the question "what produced this byte" becomes an inference from timestamps. The first is correct and potentially enormous. The second is cheap and quietly lies.
Decision
A rendition is identified by a digest over (source version, recipe digest, encoder build). An encoder upgrade therefore makes the catalogue a cache miss, which is affordable only because of ADR-01: the miss is filled lazily, by demand, and never swept.
How it is realised on AWS
rendition_digest is the primary key of the rendition table and the object key prefix in the rendition bucket. Segments are content-addressed under their rendition. A manifest references digests, so a player's request names exactly what the recipe would produce. The build-rollout view makes the consequence explicit: a promoted build is pinned by new recipes only, existing recipes and renditions are untouched, and the catalogue migrates as demand asks for renditions that do not exist yet.
Options weighed
  • ChosenContent hash over source, recipe and encoder build: Correct, traceable, and affordable only in combination with lazy regeneration. Makes an upgrade a demand-driven migration with no sweep to schedule or fund.
  • Right elsewhereAssigned identity per (title, rung, codec): Right where renditions are pre-encoded and preserved: the upgrade is a no-op and the slot semantics match the storage model. Costs the ability to say what produced a byte.
  • RejectedContent hash excluding the encoder build: Makes the upgrade free and makes two bytes with the same digest potentially different, which destroys the property the digest exists for.
  • RejectedHybrid: assigned identity with the build recorded as metadata: Traceability becomes advisory. Once a digest is not derived from the build, nothing enforces that the metadata is right.
Consequences
What it buys
  • A codec or encoder migration costs nothing to schedule: there is no sweep, no migration plan and no catalogue-wide re-encode budget.
  • Mixed-build output within a title is impossible by construction, because every rendition's digest names its build.
  • A cache is trustworthy: if the digest matches, the bytes are the bytes the recipe specifies, so no verification step is needed on a hit.
  • Regression hunting is a query — which builds produced the quarantined renditions — rather than a bisection.
What it costs
  • Promoting a build invalidates nothing and misses everything: every subsequent regeneration is a new encode, so build promotions have a diffuse compute cost that is hard to attribute.
  • Digests are opaque, so every human-facing surface needs a translation layer to say which title and rung a digest belongs to.
  • A title can hold renditions from several builds simultaneously during a long migration, which is correct and complicates quality comparisons.
  • Build images must be retained for as long as any recipe references them.
Choose differently when
If encoder builds were promoted very frequently — weekly, say — the diffuse regeneration cost could overwhelm the saving, and the right answer becomes batching promotions to a slower cadence rather than changing the identity scheme. If the design ever abandoned ADR-01 and pre-encoded everything, assigned identity becomes the better choice immediately, because a catalogue-wide cache miss with no lazy fill is a catalogue-wide re-encode.
Why it holds up over time
Content addressing has won in every adjacent domain — container images, build artefacts, package registries, version control — for the same reason it wins here. Codec generations will keep arriving, roughly one significant one every few years, and each arrival makes a lazily-migrating identity scheme more valuable than a scheme that requires a funded sweep.
LessonA naming scheme is a migration policy in disguise. Decide what you want an upgrade to cost, then pick the identity that produces that cost.
Shown on views12 16 19

encode-executionTurning a recipe into bytes on compute that is being taken away.

ADR-05

Bound reclaim loss with small chunks rather than checkpointing encoder state

Accepted

When the platform loses a worker mid-encode, does it lose one small unit of work, or does it preserve the work by serialising and reviving the encoder's state elsewhere?

Context
The economic premise of the whole design is reclaimable capacity: an assumed 18% hourly reclaim at steady state and up to 60% of the fleet withdrawn inside a ten-minute window. At that rate a whole-file encode of a feature film would rarely finish. Two mechanisms can rescue it. Splitting the source into independently encodable chunks bounds the loss to one chunk but multiplies per-task setup cost and, more importantly, multiplies the number of boundaries at which rate control restarts — which is the cause of visible quality pumping at joins. Checkpointing preserves the work but requires the encoder to serialise its internal state and another machine to revive it, which most production encoders cannot do and none do portably.
Decision
Split on shot or GOP boundaries into chunks sized so expected reclaim loss stays within a stated bound — an assumed one chunk-minute — and reschedule a lost chunk without operator action. Chunk duration is a tunable policy value, not a constant. No encoder state is serialised or revived.
How it is realised on AWS
The recipe carries the chunk plan derived from shot boundaries, so chunk edges land where a cut already breaks temporal prediction and the boundary is cheapest to make invisible. Chunk tasks are rows in Spanner claimed under a lease; lease expiry is the reclaim signal, and a requeue is one row update. Partial output in chunk scratch is discarded rather than salvaged. The swimlane view draws Reclaim as one of six columns precisely because it is a normal stage of the flow rather than an error path.
Options weighed
  • ChosenShot-boundary chunks with bounded loss and reschedule: Works with any encoder, needs no state serialisation, and puts the quality cost where a cut already is. Chunk size becomes the dial between loss and boundary count.
  • DeferredCheckpoint encoder state and resume on another worker: Preserves work and is the right answer for very long or very complex sources where chunking costs too much quality. Phase 3, and only if the chosen encoder can serialise state at all.
  • Right elsewhereWhole-file encode on reserved capacity: Right for a small volume of premium titles where boundary artefacts are unacceptable and the compute premium is affordable. Does not scale to 3,500 source hours a day.
  • RejectedFixed-duration chunks irrespective of content: Simplest to schedule and puts boundaries in the middle of motion, which is exactly where a restart of rate control is most visible.
Consequences
What it buys
  • Job progress is monotonic across any number of reclaims, so a premiere encode survives a withdrawal storm without operator involvement.
  • The mechanism is encoder-agnostic, so the design is not hostage to one encoder's feature set.
  • Chunk duration gives a single dial that trades reclaim loss against boundary count, measurable rather than argued.
  • Parallelism across thousands of chunks is what makes the assumed 45-minute premiere turnaround arithmetically possible at all.
What it costs
  • Every boundary is a potential artefact, and verifying boundaries is a mandatory structural check rather than an optional one.
  • Per-chunk setup — container start, recipe fetch, source range read — is paid thousands of times per title and is a real fraction of the bill for short chunks.
  • Rate-control continuity across chunks has to be handled by the encoding strategy, which constrains which encoder settings are usable.
  • Chunk scratch holds partial outputs with a TTL, which is storage spent on work that will sometimes be thrown away.
Choose differently when
Two measurements would flip it. If the prototype finds no chunk duration at which boundaries are perceptually invisible for grainy or high-motion content, checkpointing moves from Phase 3 to mandatory for that content class. And if reclaim intervals turned out to be much longer than assumed — a region and machine family with 2% hourly reclaim rather than 18% — larger chunks or whole-file encodes become viable and the setup overhead argument reverses.
Why it holds up over time
The decision depends on reclaimable compute being materially cheaper than committed compute, which is structural: providers price interruptible capacity to monetise idle fleet, and that gap has widened rather than narrowed. Encoders have also not converged on portable state serialisation in twenty years, so the alternative has not become easier. What will change is the chunk size, as machine types and codecs change — which is why it is a policy value.
LessonWhen the platform cannot stop losing workers, buy resilience with task granularity rather than with worker protection — and then treat the granularity as a measured dial rather than a constant someone once chose.
Shown on views14 13 22
ADR-06

Determinism is an encoder release gate, not a hoped-for property

Accepted

Does the platform require byte-identical output for the same source, recipe and build, and does it refuse to promote an encoder that cannot deliver it?

Context
Three separate properties of this architecture quietly assume determinism. Idempotent retries assume a duplicate execution is indistinguishable from a single one, which is only true if the output is identical. Content-addressed identity assumes a digest over inputs predicts the output, which is only true if the mapping is a function. And lineage assumes that recording a recipe and a build explains a published byte, which is only true if those inputs fully determine it. Encoders are full of things that break this: thread-count-dependent partitioning, time-based decisions, hardware-specific instruction paths, non-deterministic rate-control heuristics. None of them announce themselves, and the symptom — two chunks of one rendition that differ subtly — is among the hardest defects to localise.
Decision
An encoder build is promoted only if it replays a golden corpus byte-identically across the machine types the fleet runs. A build that cannot is rejected before anything else about it is measured. Quality regression against the current build is a separate, subsequent gate.
How it is realised on AWS
The rollout flow puts the determinism gate immediately after the build, before the regression gate, the shadow encode and the canary. The gate re-encodes a fixed corpus spanning grain, animation, high motion and HDR, on every machine family in the fleet, and compares digests. The build registry holds only approved digests, and a worker whose running build does not match the recipe's pinned build refuses the task rather than producing output nobody can reproduce.
Options weighed
  • ChosenDeterminism as a hard promotion gate: Protects idempotence, content addressing and lineage in one step, and catches the failure in CI rather than in a year-old rendition.
  • RejectedBest-effort determinism with equivalence checking at use: Means verifying every cache hit, which removes the cache's entire benefit, and turns every regeneration into a quality comparison.
  • DeferredPin machine family as well as build to obtain determinism: A legitimate fallback if a strongly preferred encoder is deterministic per machine family but not across them. Costs scheduling flexibility, which costs Spot availability.
  • Right elsewhereAccept non-determinism and use assigned identity instead: Coherent for a pipeline that pre-encodes and preserves renditions, where nothing depends on reproducing a byte. Not coherent with ADR-01 or ADR-04.
Consequences
What it buys
  • A duplicate chunk execution is free, so at-least-once dispatch is sufficient and no distributed lock is needed on the hot path.
  • A cache hit needs no verification: a matching digest means the bytes are the specified bytes.
  • Lineage claims are true rather than aspirational, which is what makes an audit or a regression hunt possible years later.
  • The failure is caught in a build pipeline against a fixed corpus, where it is cheap, instead of in a stitch defect report.
What it costs
  • Some otherwise excellent encoders and some hardware-accelerated paths will fail the gate and be unavailable.
  • Encoder settings that improve quality through non-deterministic heuristics are off the table, which may cost measurable bitrate.
  • The golden corpus and the cross-machine-type replay are real infrastructure to build and maintain.
  • Upstream encoder releases may be blocked for months while determinism is fixed, so the platform tracks upstream more slowly.
Choose differently when
If the best available encoder for a mandatory codec were irreducibly non-deterministic, the gate becomes unenforceable for that codec. The ordered fallbacks are: pin machine family; then accept per-codec equivalence checking with a documented loss of cache trust; and only then reconsider ADR-04. Reversing the gate wholesale without reversing ADR-01 and ADR-04 would leave three decisions resting on an assumption that is no longer true.
Why it holds up over time
Reproducible builds have moved from a research interest to a baseline expectation across packaging, container and supply-chain tooling in a decade, and the same pressure is arriving in media tooling. The requirement is more likely to become easier to satisfy than harder. What will not change is that three of this architecture's load-bearing properties are consequences of it.
LessonIf several decisions quietly depend on one property, make that property a gate rather than an assumption — and put the gate where failing it is cheap.
Shown on views16 14 12
ADR-07

Priority lanes with capacity floors, ordered by deadline and work remaining

Accepted

How does a day-and-date premiere get through a fleet that is simultaneously grinding through a catalogue backfill, without the backfill being starved forever?

Context
The workload has two shapes that meet in one fleet. A premiere is small, urgent and has an immovable embargo. A backfill is enormous, patient, and will consume every available slot if allowed to. Pure priority ordering starves the backfill indefinitely, which sounds acceptable until a rights window expires on an unencoded title. Pure fairness misses premieres. Static priority fields drift: a title marked urgent six weeks ago is not urgent today, and a title nobody flagged is on air tomorrow. Meanwhile the fleet's own capacity is varying continuously because it is reclaimable, so any scheduling decision has to survive the fleet halving under it.
Decision
Work is admitted through named lanes with guaranteed minimum capacity floors, so a premiere lane cannot be starved by a backfill lane and a backfill lane cannot be starved indefinitely. Within a lane, ordering is derived from the per-title deadline and the work remaining rather than from a static priority field. Per-owner quotas and per-title spend ceilings are enforced at admission as hard stops.
How it is realised on AWS
Lane accounting lives in Spanner beside the chunk tasks, so a claim and a lane decrement are one transaction. The scheduler runs on regional GKE and re-derives ordering continuously from deadline minus estimated remaining chunk-seconds, which means a job that is falling behind rises automatically and a job with slack falls. Floors are expressed as a share of currently available slots rather than an absolute count, so a reclaim storm shrinks every lane proportionally instead of collapsing the smallest. Progress is published as completed chunk-seconds against total so an operator sees a slip coming while there is still time to act.
Options weighed
  • ChosenNamed lanes with capacity floors, deadline-derived ordering within a lane: Bounds both failure modes, and makes urgency a computed property that cannot go stale.
  • RejectedStrict priority ordering: Simple and starves the backfill until a rights window expires on a title nobody was watching the queue for.
  • RejectedFair-share scheduling across owners: Correct for multi-tenant fairness and wrong for deadlines: it has no concept of an embargo at midnight.
  • Right elsewhereSeparate fleets per lane: Right where lanes have genuinely different hardware needs or isolation requirements. Here it fragments Spot capacity, which is the scarce resource.
Consequences
What it buys
  • A premiere and a catalogue migration coexist in one fleet with a bounded worst case for each.
  • Urgency cannot go stale, because it is derived from the deadline rather than asserted once at submission.
  • Floors expressed as a share of available slots mean a reclaim storm degrades every lane proportionally rather than destroying one.
  • Spend ceilings at admission stop a runaway recipe before the money is spent rather than alerting afterwards.
What it costs
  • The scheduler needs a remaining-work estimate per job, and a bad estimate mis-orders the queue in a way that is hard to notice.
  • Lane floors are a tuning surface that will be argued about, and a wrong floor is invisible until the month a premiere is late.
  • Deriving order continuously means the queue is not stable, which makes 'why did my job not run' a harder question to answer.
  • A hard spend stop will occasionally block legitimate work at an inconvenient moment, and that is the intended behaviour.
Choose differently when
If the fleet were large enough relative to the workload that contention disappeared, lanes become bookkeeping and strict priority would do. If the opposite — a persistently over-subscribed fleet — the floors stop being floors in any meaningful sense and the real answer is capacity, not scheduling. And if deadlines proved unreliable as data, because nobody maintains the release calendar, derived ordering is no better than a static field and lanes would carry the whole weight.
Why it holds up over time
Deadline-derived scheduling over interruptible capacity is the stable shape for any workload mixing urgent small jobs with patient large ones, and it long predates this design. The specific floors and the estimator will be re-tuned continuously; the decision to compute urgency rather than declare it is what survives.
LessonPriority that is asserted once is wrong by the second week. Derive urgency from the thing that is actually true — a deadline and the work left — and bound the starvation you are willing to accept in both directions.
Shown on views14 13 06
ADR-08

Batch and on-demand generation share code, never capacity

Accepted

Does a viewer waiting for a rendition that does not exist get served by the same fleet that is grinding through the backfill, or by a separate one?

Context
ADR-01 creates a second encoding workload with completely different properties. Batch encode is throughput-bound, tolerant of minutes of queuing, and economically dependent on reclaimable capacity. On-demand generation is latency-bound with an assumed p95 of 1.5 seconds to first byte, and a viewer is sitting in front of it. Running both on one fleet means one of two bad outcomes: either the waiting viewer joins a Spot queue behind 200,000 backfill chunks, or the batch fleet inherits an availability and latency requirement that destroys the Spot economics that justified it.
Decision
The two paths share the encoder, the recipe and the code, and share no capacity. Batch runs on reclaimable Spot managed instance groups with no availability target. On-demand runs on a warm, reserved GKE pool beside the packaging origin, with request coalescing on the rendition digest, a spend ceiling of an assumed 6% of total encode spend, and a nearest-resident-rung fallback when the budget cannot be met.
How it is realised on AWS
The deployment view separates them physically: zonal Spot MIGs for batch, a regional warm pool for on-demand, with the origin and the manifest cache in the same latency-bound tier. Both consume the same recipe by digest, which is what makes the shared-code claim true rather than nominal — an on-demand regeneration produces the same digest the batch fleet would have produced. Coalescing happens at the origin, so a hundred concurrent requests for a missing rendition produce one encode. The budget check sits at admission to the warm pool, and exceeding it serves a fallback rather than queuing.
Options weighed
  • ChosenShared code, separate capacity, separate SLOs: Each workload gets the capacity model its requirement implies, and determinism guarantees the outputs are interchangeable.
  • RejectedOne fleet with priority preemption for on-demand work: Preempting a Spot worker that may itself be reclaimed gives the viewer a queue with two sources of delay and no bound.
  • DeferredOn-demand generation at the edge: Attractive for first-byte latency and currently incompatible with master access and key handling. Question 7 keeps it open.
  • Right elsewhereNo on-demand path; pre-encode everything: Right under a pre-encode-and-preserve model. Reverses ADR-01, and with it the storage and DR economics.
Consequences
What it buys
  • A waiting viewer never queues behind batch work, and the batch fleet keeps its Spot pricing and its absence of an availability target.
  • Determinism makes the two paths' outputs identical, so a rendition's provenance does not depend on which fleet produced it.
  • The on-demand path has its own spend ceiling, which turns an eviction-policy mistake into a measurable signal instead of a surprise invoice.
  • Coalescing bounds the cost of a sudden spike on a cold title to a single encode.
What it costs
  • Reserved warm capacity is paid for whether or not it is used, and sizing it is a forecasting problem with no good data on day one.
  • Two fleets mean two sets of operational behaviour, two scaling policies and two things to be on call for.
  • The fallback contract — serve a lower rung rather than wait — has to be acceptable to the product, and that is a conversation, not a configuration.
  • Archive-class master restore sits inside the first-byte budget, which couples a latency SLO to a storage class decision.
Choose differently when
If on-demand demand were low and bursty enough that a serverless cold start fit inside the budget, the reserved pool becomes waste and the right answer is scale-to-zero. If it were high and steady, the pool stops being an exception and the honest conclusion is that eviction is too aggressive — which is why the JIT share of spend is an operational metric rather than a finance line.
Why it holds up over time
Separating latency-bound from throughput-bound work onto different capacity is one of the most stable patterns in systems design, and nothing about media encoding makes it less true. What will move is the boundary: as cold-start times fall and edge compute matures, the warm pool may become serverless or move outward, and both are changes of realisation rather than of decision.
LessonTwo workloads that share code are not one workload. Let them share the code and give each the capacity model its own requirement implies.
Shown on views15 17 08

quality-and-publicationProving an output is good enough, and putting it in front of people.

ADR-09

The quality gate blocks publication, and a waiver is a recorded act

Accepted

Can a rendition that scores below the quality floor be published, and if so by whom and with what record?

Context
At an assumed 1,800 title versions a day, nobody watches the output. Quality has to be decided by a number or it is not decided at all. But a number set by a quality engineer will occasionally refuse a title that content operations needs on air tonight, and the organisation will resolve that conflict somehow — either through a documented mechanism or through a quiet configuration change that nobody can find afterwards. The design's only real choice is whether the override exists in the open.
Decision
The gate is blocking by default: a rendition below the configured floor is not publishable. The only route past it is a recorded waiver by an authorised approver, with actor and justification in the append-only audit log. The person who sets the floor and the person who may waive it are deliberately different roles.
How it is realised on AWS
The gate decision service combines a full-reference perceptual score per rung — aggregate and worst-window, against assumed floors of 93 at the top rung, 78 at the lowest, and no two-second window below 60 — with structural checks for sync drift, loudness, black and frozen frames, missing segments and decodability across the device matrix. Failure produces a diagnosis naming rungs, windows and checks, and a proposed recipe adjustment; the retry runs under a new recipe version rather than the same one, because re-running a deterministic encode would fail identically. Waivers are rows in the audit log, and the human sampling queue is weighted towards narrow passes.
Options weighed
  • ChosenBlocking gate, recorded waivers, separated roles: Makes the conflict visible and auditable rather than resolving it through an untracked configuration change.
  • RejectedAdvisory gate with a quality report: At this volume an advisory signal is an unread signal. Nothing would be refused and the floor would be decorative.
  • RejectedBlocking gate with no override at all: Sounds rigorous and guarantees that the first genuinely urgent exception is handled by someone editing the floor, which is worse than a waiver.
  • Right elsewhereHuman review of every title: Right for a premium catalogue of a few hundred titles a year. Does not survive 1,800 versions a day.
Consequences
What it buys
  • A quality failure is a blocked publish rather than a customer complaint, which is the cheapest place to find it.
  • Overrides are data: the audit log shows whether waivers cluster around one supplier, one codec or one deadline, which is how a floor gets corrected.
  • Separating who sets the floor from who may waive it keeps deadline pressure away from the bar itself.
  • A diagnosis plus a proposed recipe adjustment makes a failure actionable rather than a dead end.
What it costs
  • A premiere can be blocked at an inconvenient hour, and the escalation path has to work at that hour.
  • The floor is a judgement call that will be wrong in both directions until telemetry corrects it, and being wrong strictly is the more visible error.
  • Perceptual scoring every rung of every title is a real compute cost on top of the encode itself.
  • A waiver culture can develop, and the only defence is that waivers are counted and visible.
Choose differently when
If post-publication telemetry showed the floor refusing renditions that viewers demonstrably do not notice, the floor should move — that is the gate working, not failing. If waivers became routine rather than exceptional, the conclusion is that the floor is mis-set or the ladder solver is under-performing, not that the gate should be advisory.
Why it holds up over time
Perceptual metrics will improve and the specific floors will move with every codec generation, so the numbers in this record are the least durable part of it. The structure — a blocking numeric gate, an explicit audited override, and a separation between the person under deadline pressure and the person who sets the bar — is organisational rather than technical and should outlast several metrics.
LessonEvery quality bar will be overridden eventually. Design the override before the first incident, name who may use it, and count it — otherwise the bar is edited instead.
Shown on views18 22 13
ADR-10

Playback telemetry is a second, post-publication gate

Accepted

Is quality assurance finished at the moment of publication, or does the platform keep judging a rendition after viewers have it?

Context
A perceptual score compares an encode to its source. It does not know that one television's decoder mishandles a particular profile, that a chipset drops frames above a certain bitrate, or that a rung is being selected far more often than the ladder anticipated. Those failures are only visible from the field, and they are the ones viewers actually experience. A pipeline whose quality story ends at publication finds out about them through customer support, by which point the title has been live for days.
Decision
Rebuffer rate, rung distribution and decode errors by device class are consumed as an inbound interface and act as a post-hoc gate that can automatically quarantine a published rendition and trigger a manifest revision. The pre-publication gate remains necessary and is explicitly not sufficient.
How it is realised on AWS
Telemetry is drawn on the integration view as inbound, beside the delivery interfaces, rather than as an export — that placement is the decision. Events land in BigQuery partitioned by publish date and joined to rendition digests, so a regression can be attributed to an encoder build or a ladder change rather than to a title. A quarantine removes the rendition from the manifests that advertise it, which is a revision rather than a withdrawal of the title; because manifests are cached for an assumed 30 seconds and segments for 30 days, the revision takes effect almost immediately and invalidates no cached byte.
Options weighed
  • ChosenTelemetry as an automatic post-publication gate: Catches the device-specific and selection-pattern failures that no full-reference metric can see, and does it in hours rather than through support tickets.
  • RejectedTelemetry as a dashboard for the quality team: Turns a gate into a report. At 1,800 versions a day a report is read about the titles someone already suspected.
  • DeferredExpand the pre-publication device matrix instead: Worth doing and cannot be complete: the matrix tests decodability on devices we have, not selection behaviour on networks we do not.
  • RejectedManual withdrawal on support escalation: The current state of the art in many pipelines, and it measures quality in units of complaints.
Consequences
What it buys
  • Device-specific failures are found by the platform rather than by viewers, and attributed to a build or a ladder rather than to a title.
  • A quarantine is a manifest revision, so remediation costs nothing at the edge and needs no re-encode to take effect.
  • The assumed false-pass target of 0.2% of titles per quarter becomes measurable, which makes the pre-publication floor tunable against evidence.
  • The same telemetry feeds ladder tuning, so the quality loop and the cost loop share one signal.
What it costs
  • Automatic quarantine can remove a rung from a working title on a noisy signal, so the thresholds need hysteresis and a blast radius.
  • Telemetry arrives with a delay and at a volume that costs real money to retain — an assumed 13 months at full granularity.
  • Attribution depends on the digest chain being intact end to end, so this gate is only as good as ADR-03 and ADR-04.
  • Device-class taxonomies drift as the player estate changes, and a stale taxonomy silently misattributes failures.
Choose differently when
If the player estate were narrow and stable — a single device family, say — a pre-publication matrix could plausibly be complete and the post-hoc gate would be redundant. If telemetry quality were poor enough that quarantines were mostly false positives, the gate should degrade to alerting a human rather than acting, which is a change of autonomy rather than of principle.
Why it holds up over time
The player estate gets more diverse over time, not less, and every new codec adds device-specific decode behaviour that no laboratory matrix fully predicts. The case for judging quality in the field strengthens as the estate fragments. What will change is how much autonomy the gate is given.
LessonA metric that compares output to input cannot see the thing that only happens on someone's television. Close the loop from the field, and give the loop the authority to act.
Shown on views09 18 22
ADR-11

Atomic manifest swap is the only publication primitive

Accepted

Does a title become available rendition by rendition as encoding finishes, or in one indivisible step?

Context
A full ladder across several codecs completes over hours. The natural implementation reveals renditions as they land, because each one is useful the moment it exists. The consequence is that for most of those hours the title is in a state nobody designed: a player may see a top rung with no fallback, or a ladder missing the rungs a phone needs, and the set of things a viewer can experience is the set of partial completions. At midnight on a premiere, that state is what the audience gets.
Decision
Publication is an atomic swap of a manifest pointer per title and territory. An incomplete ladder publishes as a narrower manifest or not at all. A launch subset may be published ahead of the full ladder, with the remainder added by later manifest revisions. Rollback is a pointer move to the previous revision, completing within an assumed 120 seconds and with no dependency on the encode fleet.
How it is realised on AWS
The publisher verifies that every segment a candidate manifest references is durably present and gate-passed before emitting it; a manifest pointing at an absent segment is treated as a publication defect rather than a cache miss. The swap updates the catalogue projection, which the origin and the catalogue service read. Because manifests are cached for an assumed 30 seconds and segments are immutable and content-addressed, a revision — adding rungs, narrowing a ladder, quarantining a rendition, withdrawing a title — propagates in seconds and invalidates nothing at the edge.
Options weighed
  • ChosenAtomic manifest swap, launch subset, revisions: Every state a viewer can reach is a state somebody chose, and rollback needs no fleet.
  • RejectedProgressive reveal as renditions complete: Cheap and makes the set of viewer-visible states equal to the set of partial completions, which nobody reviewed.
  • RejectedPublish only when the full ladder is complete: Safe and misses the premiere: the full ladder takes an assumed four hours and the embargo is at midnight.
  • RejectedPer-rendition availability flags read by the player: Moves the publication decision into the player, which is out of this boundary and cannot be rolled back from here.
Consequences
What it buys
  • Nothing is ever half-published, and every reachable state is a deliberate one.
  • Rollback is a pointer move with no encode, no packaging and no re-upload, which is what makes the assumed 120-second target credible.
  • A launch subset lets a premiere ship on time without pretending the full ladder is ready.
  • Immutable segments plus a 30-second manifest cache mean revisions are near-instant and free at the edge.
What it costs
  • The publisher must verify durability across potentially thousands of segments before each swap, which is latency on the critical path at midnight.
  • Launch-subset width is a judgement call with a real product consequence, and on day zero there is no demand signal to inform it.
  • A narrow first manifest means some viewers genuinely get fewer rungs for a while, and that has to be acceptable.
  • Manifest revisions multiply per title and territory and device class, so the projection carries more rows than an intuitive design would.
Choose differently when
If the full ladder could be produced inside the embargo window — a much faster fleet, or a much narrower ladder — the launch subset becomes unnecessary complexity and publishing complete is simpler and better. If territories and device classes multiplied far beyond the assumption, per-title atomicity might need to become per-title-per-territory batching to keep the swap bounded, which is a change of granularity rather than of principle.
Why it holds up over time
Atomic pointer swaps over immutable content is the same pattern as a blue-green deploy, a database index swap and a content-addressed release: it has been the right answer in every domain that needed a reversible publication, and it does not depend on anything about video. Segment immutability is what makes it cheap, and immutability is not going away.
LessonIf a long-running process can produce many intermediate states, publish an indivisible pointer instead — and the states nobody chose become unreachable rather than merely unlikely.
Shown on views13 19 04

economics-and-lifecycleDeciding which renditions are worth existing, and for how long.

ADR-12

The bitrate ladder is solved per title from a complexity profile

Accepted

Does every title get the same ladder, a ladder solved for that title, or a ladder that varies shot by shot within the title?

Context
A fixed ladder is wrong for almost everything. Animation at 1080p needs a fraction of the bitrate of grainy 35 mm film at the same resolution and perceptual quality; sport needs more than either. A fixed ladder therefore wastes bitrate on the easy content and starves the hard content, and does both at every rung. Per-title solving corrects that at the cost of analysis compute. Per-shot solving corrects it further and introduces a harder problem: rungs must stay segment-aligned across a title so a player can switch at any segment, and per-shot variation pulls against that.
Decision
The ladder is derived per title from a complexity profile — shot boundaries, per-shot spatial and temporal complexity, grain, dominant motion, letterbox geometry, and the effective source resolution after any upscale in the master. The egress saved is weighed against the analysis compute spent, per title, and the comparison is recorded. Per-shot variation is deferred to Phase 3, with segment alignment preserved as a hard constraint.
How it is realised on AWS
The analyser runs as Cloud Run jobs before any encode is scheduled, and the solver writes its result into the recipe, so the ladder is frozen before a single chunk runs. Rungs above the effective source resolution are omitted rather than upscaled, and a rung exists only if some supported device class and some observed network condition can select it. The cost comparison is a row in BigQuery per title, which is what makes the trade arguable rather than asserted: at an assumed 40 PB/day of delivered egress, one percent of average bitrate across the ladder is worth about 400 TB/day.
Options weighed
  • ChosenPer-title ladder from a complexity profile: Captures most of the available bitrate saving at a bounded analysis cost, and keeps segment alignment trivially intact.
  • RejectedFixed catalogue-wide ladder: Simplest, cheapest to operate, and wrong in both directions simultaneously on a catalogue spanning animation and 35 mm grain.
  • DeferredPer-shot ladder variation: More saving and a real risk to cross-rung segment alignment. Phase 3, and only once the per-title gain has been measured.
  • RejectedPer-device-class ladders: Multiplies renditions by device class, which multiplies the storage and cache problem ADR-01 exists to contain.
Consequences
What it buys
  • Easy content stops being over-encoded and hard content stops being starved, at every rung rather than on average.
  • The ladder decision becomes answerable in money, which is what lets it be tuned rather than debated.
  • Refusing to upscale removes a whole class of rungs that cost storage and egress and deliver nothing.
  • Freezing the ladder into the recipe before any chunk runs means the whole title is encoded against one coherent decision.
What it costs
  • Analysis is compute spent before any output exists, on every title, including the ones nobody will watch.
  • Variable ladders make cross-title quality comparison harder, because two titles' rung-three are no longer the same thing.
  • The solver is a model, and a model that drifts produces ladders nobody notices are wrong until telemetry says so.
  • Per-title rung counts make cache and storage forecasting less predictable than a fixed ladder.
Choose differently when
If analysis cost per source hour approached the encode cost — a far more expensive analysis, or much cheaper encoding — the per-title gain would stop paying for itself on the long tail, and the right answer becomes per-category ladders with per-title solving reserved for high-demand titles. Conversely, if the measured saving were large and segment alignment proved tractable under per-shot variation, Phase 3 moves forward.
Why it holds up over time
Content-adaptive encoding has moved from a research result to standard practice over the last decade, and every codec generation widens the gap between a fixed ladder and a solved one because codecs get better at exploiting content structure. The specific profile features and the solver will be replaced repeatedly; deciding per title rather than per catalogue is what survives.
LessonCatalogue-level defaults are where the waste lives. Make the expensive decision once per item, record what it cost to make, and the trade stops being a matter of opinion.
Shown on views02 10 18
ADR-13

Rendition residency is an economic decision, re-made continuously

Accepted

On what rule is a rendition demoted to a cheaper storage class, and on what rule is it deleted outright?

Context
ADR-01 permits eviction; it does not say when. The obvious rule is age, because every object store implements it and it needs no thought. Age is also almost unrelated to the thing that matters: a ten-year-old title in a prestige collection may be watched steadily while a title from last month is watched twice. The quantity that actually decides whether a rendition should exist is the cost of keeping it against the cost of making it again, over whatever horizon the platform is willing to plan for — and both sides of that comparison move with storage prices, compute prices and the title's own demand.
Decision
A rendition is demoted and eventually evicted when its regeneration cost falls below its retention cost over the policy horizon. The rule is economic, not temporal, and is re-evaluated continuously against observed demand. Eviction is not an end state: a request for an evicted digest re-enters the lifecycle at generation.
How it is realised on AWS
A demand-tier controller on GKE reads per-digest request rates from the telemetry store and writes a residency class onto each rendition row; Cloud Storage lifecycle rules then execute the demotion and deletion that the controller has decided. The lifecycle view is drawn as a ring with seven states precisely so that Evicted has an outgoing arrow — a request re-enters at Generated through the on-demand path of ADR-08. Targets are assumed: at least 70% of catalogue hours held at the cheapest tier or not held at all, rendition storage no more than 18% of total storage spend, and on-demand generation no more than 6% of encode spend.
Options weighed
  • ChosenEconomic rule on demand, regeneration cost and retention cost: Compares the two quantities that actually decide the question, and re-prices itself as costs and demand move.
  • RejectedAge-based lifecycle rules only: Free to implement and uncorrelated with demand. Evicts the steadily-watched prestige title and keeps last month's failure.
  • Right elsewhereManual curation of residency by content operations: Right for a few hundred titles with strong editorial opinions. Unworkable across 1.2 million rendition sets.
  • RejectedKeep everything resident in the cheapest class: Avoids every fallback and keeps the full object count, which is a bill that scales with catalogue rather than viewing.
Consequences
What it buys
  • Residency re-prices itself as storage and compute costs move, without anyone rewriting a policy.
  • The long tail costs roughly what it is worth, which is the saving ADR-01 was taken for.
  • The JIT share of spend becomes a single number that says whether the policy is tuned, visible on the observability grid.
  • A sudden spike on a cold title promotes it automatically, so popularity fixes its own residency.
What it costs
  • Aggressive eviction converts into viewer-visible fallbacks, and the feedback loop between the two is indirect and delayed.
  • The controller needs per-digest demand data at a granularity that is itself expensive to collect and retain.
  • Regeneration cost is an estimate, and a wrong estimate evicts things it should not in a way that is hard to notice.
  • Two systems now decide residency — the controller and the lifecycle rules — and a disagreement between them is a confusing class of bug.
Choose differently when
If the on-demand spend ceiling were breached persistently, the policy is too aggressive and the horizon should lengthen rather than the ceiling rise. If a rights holder's contract forbade disposing of deliverables for part of the catalogue, residency becomes a per-title contractual floor, which this design can express as a policy value — which is why it is a policy value and not an architectural constant.
Why it holds up over time
The comparison itself — keep versus remake — is timeless, and expressing it as a rule rather than a schedule means the design absorbs changes in both prices without redesign. What will change continuously is the horizon and the thresholds, which is the intended behaviour rather than a weakness.
LessonWhen the obvious policy lever is time and the real variable is money, write the rule in money. Age-based retention is a proxy nobody chose and everybody inherits.
Shown on views19 11 15

protection-and-operationsKeeping an unreleased master and a content key safe, and recovering a region.

ADR-14

Regional recovery is regeneration from custody, not rendition replication

Accepted

After the loss of the encode region, is the platform recovered by replicating renditions into a second region, or by regenerating them there from masters and recipes?

Context
This is ADR-01 meeting a disaster-recovery review, and it is where the cache model is most likely to be quietly reversed. The familiar answer is to replicate everything to a standby region and fail over, because that is what the runbook template says. Applied here it means replicating an assumed 1.2 million rendition sets, paying twice for storage and once for egress on everything, in order to protect bytes that the architecture has already established are a function of two much smaller things.
Decision
Masters, recipes, gate scores and the audit log are dual-region. Renditions are not replicated for availability — they are regenerated. The standby region carries a warm control plane and no encode fleet. After a region loss, the published manifest degrades to the surviving rung set and refills, with an assumed 30-minute p95 for a title's launch subset.
How it is realised on AWS
The deployment view shows three groupings: a primary encode region with the fleet, a dual-region custody band holding masters, recipes and audit, and a standby region with a warm control plane and a degraded origin and nothing else. Control-plane recovery is an assumed RTO of 15 minutes against an RPO of 5 seconds on Spanner. In-flight jobs restart from completed chunks, which is a direct consequence of chunk idempotence in ADR-05. Cross-region regeneration is exercised on a schedule in Phase 3, because a recovery mechanism nobody drills is a hypothesis.
Options weighed
  • ChosenDual-region custody, regeneration on failover, no standby fleet: Replicates the small irreplaceable things and buys a regeneration budget instead of a second copy of everything.
  • Right elsewhereFull rendition replication to a standby region: Right under a pre-encode-and-preserve model, or where a contract specifies a warm second copy. Doubles storage to protect derivable bytes.
  • RejectedActive-active encode in both regions: Doubles the fleet footprint and adds a cross-region consistency problem on job state that nothing in the requirement asks for.
  • RejectedBackup and restore of renditions from archive: Slower than regenerating them, and pays archive storage for every rendition to do it.
Consequences
What it buys
  • Disaster recovery costs the replication of masters and recipes — orders of magnitude smaller — plus a regeneration budget.
  • The standby region is cheap enough to keep genuinely warm, because it holds a control plane rather than a fleet.
  • Chunk idempotence means in-flight jobs resume rather than restart, so a failover does not lose a night's encoding.
  • The degraded state is explicit and drawn: a narrower rung set, refilling, rather than an outage.
What it costs
  • Immediately after a failover, viewers in every territory may get a narrower ladder than usual until the refill completes.
  • Recovery time now depends on encode capacity in the standby region, which is the resource that was deliberately not pre-provisioned.
  • A regeneration storm during a failover competes with the on-demand path for exactly the capacity that is scarcest.
  • The mechanism is only credible if it is drilled, which is ongoing operational cost rather than a one-off build.
Choose differently when
If a contract or a regulator required a warm second copy of deliverables rather than a demonstrated ability to reproduce them, replication becomes mandatory for that part of the catalogue. If regeneration throughput in a standby region proved unobtainable at short notice — a capacity constraint rather than a cost one — the honest fallback is replicating the launch subset only, which protects the viewer-visible case at a fraction of the full cost.
Why it holds up over time
The decision follows from ADR-01 and inherits its durability. The specific regional topology will change; the principle that you replicate the inputs to a function rather than its outputs is what persists, and it gets more valuable as the ratio between catalogue size and master size grows.
LessonReplicate the inputs, not the outputs. A disaster-recovery plan that copies everything derivable is usually a plan written before anyone asked what was derivable.
Shown on views17 11 22
ADR-15

The encode fleet is an untrusted zone, and no fallback weakens protection

Accepted

What is the assumed attack on this platform, and what is the one thing the design refuses to degrade under any failure?

Context
Most platforms place their trust boundary at the API and treat internal compute as trusted. That is the wrong boundary here. The encode fleet's entire job is to run a decoder over bytes supplied by third parties, and codec parsers are among the most consistently exploitable software in production anywhere. Meanwhile the assets involved are unusually sensitive in a specific way: an unreleased master is a pre-embargo copy of something whose leak is a contractual and commercial event, and a content key is the thing that makes an entire catalogue's encryption meaningful. Both of those are handled by the same fleet.
Decision
The encode fleet is treated as untrusted. Workers decode inside a sandbox with no egress, hold a short-lived workload identity granting read on one master prefix and write on one rendition prefix, and nothing else. Content keys are obtained from an external key service per packaging job, held in memory for that job, and never persisted in pipeline stores, logs, metrics or worker disk. And one invariant admits no exception: no fallback may weaken protection — quality, latency and rung availability may all degrade, encryption may not.
How it is realised on AWS
The trust-zone view gives the fleet its own zone between the perimeter and the trusted platform, which is unusual and deliberate. gVisor isolates the decoder process; egress is denied at the network level so a compromised parser has nowhere to send anything. Pre-release masters sit in the custody zone with CMEK, per-object read auditing and alertable bulk reads, with per-rights-holder key separation where a contract requires it, and pre-release renditions are served from an origin that is not publicly routable until publication. The invariant shows up concretely in the failure table: a key service outage stalls publication and that is the whole of the handling — there is no unencrypted path to fall back to.
Options weighed
  • ChosenUntrusted fleet zone, sandboxed decoders, keys never persisted, protection never degraded: Places the boundary where the hostile input actually arrives, and removes the one fallback that would be reached for under pressure.
  • RejectedTrusted internal compute with a perimeter boundary: The conventional topology, and it puts the trust boundary in the one place the attack does not come from.
  • RejectedCache content keys in the packaging tier for throughput: Trades a contractual exposure for a latency gain nobody asked for, and makes the pipeline a place keys live.
  • RejectedAllow an unencrypted path for internal review workflows: Creates exactly the exception that gets used during an incident, and it would be used by someone with a deadline rather than a threat model.
Consequences
What it buys
  • A codec-parser compromise is contained to a sandbox with no egress and credentials that reach two prefixes.
  • Key custody never passes to the pipeline, so a pipeline compromise does not become a catalogue-wide decryption event.
  • Per-object read auditing on masters makes a pre-embargo leak investigable rather than deniable.
  • Removing the unencrypted fallback removes the option that would otherwise be taken at 2am under commercial pressure.
What it costs
  • Sandboxing costs encode throughput, and on a fleet of this size that is a measurable share of the compute bill.
  • Per-job key fetches put the key service on the packaging critical path, so its availability becomes a publication dependency.
  • A key service outage delays a premiere, and the design offers no mitigation beyond retry — by intent.
  • Narrow per-worker credentials and no egress make debugging a failing encode genuinely harder for operators.
Choose differently when
Nothing in ordinary operation flips the protection invariant; it exists precisely to be unflippable. The realisation can change: if a provider offered hardware-level isolation for the decode step with less overhead than a sandbox, that replaces gVisor without touching the decision. If key-service availability proved to be the dominant cause of missed premieres, the right response is a more available key service, not a weaker encryption path.
Why it holds up over time
Codec parsers have been a reliable source of memory-safety vulnerabilities for as long as codecs have existed, and nothing about the current direction of either media tooling or sandboxing changes that. The sensitivity of pre-release content is contractual rather than technical and does not decay. Sandbox overheads are falling, which makes the decision cheaper over time.
LessonPut the trust boundary where the hostile input arrives, not where the authentication is. Then name the one thing you will never degrade, and delete the fallback that would let you.
Shown on views20 21 22

Every package used, in one table

Ten terms that carry specific meaning in this package, several of which mean something looser in general industry use.

PackageWhat it isWhat it does hereConsidered instead
Mezzanine The high-bitrate intermediate master delivered by a rights holder — an assumed 80 Mbps, about 36 GB per source hour — from which every rendition is derived. The only non-derivable artefact in the pipeline. Immutable on arrival, retention-locked, dual-region. Sometimes called the source, the house master or the ProRes; in this package it is always the delivered intermediate.
Recipe An immutable, content-addressed document holding the solved ladder, the analysis inputs, every encoder parameter and the pinned encoder build. The authoritative derived artefact and the only sanctioned encode input. The join key between a master and everything derived from it. Often called a job template or an encoding profile, both of which usually imply something mutable and shared.
Rendition One encoded version of a title at one rung, in one codec, produced by one recipe and one encoder build. A cache entry: evictable under an economic policy, regenerable on demand, carrying no RPO. Elsewhere a rendition is typically a durable deliverable. That reading is explicitly rejected here (ADR-01).
Rung One step of the bitrate ladder — a resolution and target bitrate pair that a player may select. The unit the ladder solver decides, the gate scores and the manifest advertises per device class. Also called a variant, a representation or a profile depending on the streaming format.
Launch subset The few rungs that carry the majority of first-day sessions, published ahead of the full ladder. What makes an assumed 45-minute premiere turnaround possible without pretending the full ladder exists. No settled industry term; sometimes described as a fast-publish or day-one ladder.
Reclaim The provider withdrawing a Spot worker mid-task — an assumed 18% hourly at steady state. The expected steady state rather than a failure. Costs at most one chunk and is drawn as a stage of the flow. Also preemption or interruption. Treating it as an error path is the mistake this design is built to avoid.
Common encryption Encrypting segments once under a scheme that multiple DRM systems can all issue licences against. What lets one encrypted segment set serve three DRM systems, so adding one is an integration rather than a re-packaging job. Sometimes conflated with DRM itself; here DRM is the external licence service and this is the encryption scheme.
Manifest The per-device-class document advertising which rungs, codecs, audio tracks and encryption schemes are available. The only mutable published artefact, cached for an assumed 30 seconds. Publication, rollback, revision and quarantine are all manifest operations. Also the playlist or the MPD. Not to be confused with the delivery manifest a supplier sends with a master.
Gate The blocking combination of a full-reference perceptual score per rung, structural checks and a device decoder matrix. Decides publishability. Passing is required; a waiver is the only route past it and is recorded with actor and reason. Often an advisory QC report. The distinction between blocking and advisory is the whole of ADR-09.
Residency Which storage class a rendition currently occupies, or whether it exists at all. Decided continuously by comparing regeneration cost against retention cost over the policy horizon, not by age. Usually expressed as a retention schedule, which is a proxy for demand that nobody actually chose.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.