Streaming Video Encoding & Packaging Pipeline

Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records

Almost nobody outside the industry has a name for this system, and almost everybody has met it. A show lands at midnight and plays instantly in 4K on the television, while the same episode on a phone starts soft and sharpens eight seconds in. A film from 1974 looks worse than it should on a big screen; a cartoon from last week looks flawless at a quarter of the bitrate. A new release stalls for everyone in the first hour and is fine by morning. None of that is the network and none of it is the player — it is the encoding and packaging pipeline, the thing that turns one enormous studio master into the few hundred small files a player actually downloads, and decides per title how many of them are worth making. This package is that pipeline for a consumer video-on-demand service: an assumed 400,000 titles and 250,000 source hours, 3,500 source hours a day arriving as ~126 TB of immutable masters, 220,000 chunk-encode tasks a day across a reclaimable fleet of ~14,000 slots, and 40 PB a day of delivered egress — built on dual-region archive custody, open-source encoders on per-second-billed Spot capacity, a transactional control plane, and a warm fleet that regenerates a rendition while a viewer waits.

22 views 22 HTML views22 SVG22 draw.io 2 documents Updated 2026-10-11
Architecture views

22 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    What the pipeline owns, and the four systems it depends on without controlling.

  2. 02
    High-Level Architecture

    Seven stages from delivery to publication, and the one loop in the whole design.

  3. 03
    Actors and Their Core Journeys

    Eight actors, their goals in their own words, and which journeys get their own map.

  4. 04
    Journey — Premiere Night

    The journey the whole launch-subset mechanism exists for, and the one phase where it fails.

  5. 05
    Journey — A Cold-Stored Title

    The price of treating renditions as a cache, paid in the place where it is cheapest.

  6. 06
    Journey — Getting a Title to Air

    The operator's journey, where the worst moment is three days before anyone presses play.

  7. 07
    Layered Architecture

    Eight layers, with the recipe plane highlighted because it holds the authoritative artefact.

  8. 08
    Platform Components

    The container view on Google Cloud, grouped by plane rather than by service type.

  9. 09
    Integration Surface

    Five inbound interfaces, four outbound, and the one that runs the wrong way on purpose.

  10. 10
    Data Flow — One Master Through

    Seven columns, of which only the first three cannot be rebuilt.

  11. 11
    Storage Zones

    Three zones ordered by what loss costs, and only the top one is backed up.

  12. 12
    Data Model

    Eleven entities, and the one key the whole model turns on.

  13. 13
    Critical Flow — Master to Manifest

    Seventeen messages, two of which are failures, and both of those are normal.

  14. 14
    Chunked Encode and Reclaim

    Six stages across five lanes, with reclaim as a column rather than an exception.

  15. 15
    On-Demand Regeneration

    What happens when a viewer asks for a rendition that does not exist.

  16. 16
    Encoder Build Rollout

    How a new encoder reaches production without re-encoding 250,000 hours.

  17. 17
    Deployment Architecture

    One encode region, dual-region custody, and a standby with no fleet in it.

  18. 18
    Observability

    Six signal types across six stages, with cost as a signal row rather than a monthly report.

  19. 19
    Rendition Lifecycle

    Seven states, and the reason the ring actually closes.

  20. 20
    Security and Trust Zones

    Five zones, and why the encode fleet gets one of its own.

  21. 21
    Identity, Keys and Licence Flow

    Thirteen messages in which the pipeline never ends up holding a key.

  22. 22
    Failure Classes

    Ten assumed failures, how each is detected, and what each one still leaves behind.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 15 min read

Architecture One-Pager

Streaming Video Encoding & Packaging Pipeline · Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records

The master plus the recipe is the asset. Every rendition is a cache entry. Everything else in this architecture is a consequence of that sentence.

Almost nobody outside the industry has a name for this system, and almost everybody has met it. A show lands at midnight and plays instantly in 4K on the television, while the same episode on a phone starts soft and sharpens eight seconds in. A film from 1974 looks worse than it should on a big screen; a cartoon from last week looks flawless at a quarter of the bitrate. A new release stalls for everyone in the first hour and is fine by morning. None of that is the network and none of it is the player. It is the encoding and packaging pipeline, and the reason it is hard is not video compression. It is that the catalogue is enormous and the viewing is not: an assumed 400,000 titles and 250,000 source hours, of which an assumed 72% of catalogue hours attract under 0.1% of total watch time. Pre-encoding a full ladder of a dozen renditions in several codecs for every one of them is a storage bill paid, in perpetuity, for files nobody requests. Not pre-encoding them means a real viewer waits. The architecture is the answer to where that line goes.

Seven stages, two execution fleets and three storage zones. A rights holder delivers a mezzanine master; the platform verifies it against a delivery manifest, probes it, classifies it against a conformance policy, and lands it write-once in dual-region archive storage. It then analyses the content — shots, spatial and temporal complexity, grain, motion, effective resolution after any upscale in the master — and solves a bitrate ladder for that title, recording the ladder, every encoder parameter and the encoder build identifier as a single immutable, content-addressed recipe. The recipe, not the rendition, is the authoritative derived artefact. A control plane on Spanner materialises the recipe into chunk tasks on shot boundaries and dispatches them to a Spot fleet whose workers are expected to be withdrawn continuously; a reclaim costs one chunk and is rescheduled without an operator. Stitched renditions pass a blocking gate — a full-reference perceptual score per rung, structural checks, and decodability on a device matrix — before a packager segments and encrypts them once under common encryption, with content keys held in memory for the duration of one packaging job and never persisted. Publication is an atomic swap of a manifest pointer, launch subset first. After publication, a lifecycle controller demotes and evicts renditions on an economic rule, and a separate warm fleet regenerates an evicted rendition when a viewer asks for it, falling back a rung rather than making anyone wait past the budget.

What it is, and what it is not

  • A pipeline that makes a master playable on every supported device — not A video compression research project
  • A system whose authoritative artefacts are a master and a recipe — not A rendition library that must be preserved
  • An economic argument about which renditions are worth existing — not A quality-at-any-cost encoder farm
  • A platform that proves quality with a number before publishing — not A workflow that relies on someone watching the output
  • Designed around compute being taken away continuously — not A highly available encode fleet
  • A publisher of immutable segments and mutable manifests — not A CDN or an edge cache
  • A consumer of a DRM licence service — not A DRM or licence-issuing system
  • Scoped to video-on-demand — not A live or contribution encoding path
  • A multiplexer of audio and subtitle tracks delivered to it — not A subtitle authoring or translation tool

The decisions that are the architecture

  1. The master and the recipe are the only assets (ADR-01) — Renditions get no backup, no cross-region replication and no RPO. Storage cost becomes a retention-versus-regeneration calculation per title, and disaster recovery becomes replication of the two small things plus a regeneration budget.
  2. The mezzanine is immutable on arrival (ADR-02) — Write-once, retention-locked, never modified in place. A redelivery is a new source version requiring explicit promotion, not an overwrite — so a supplier cannot silently change what was published.
  3. The recipe is the only sanctioned encode input (ADR-03) — Immutable, content-addressed, pinned to an encoder build. It is the join key between a master and everything derived from it, and the reason an encode can be re-run identically in eighteen months.
  4. Renditions are identified by content hash (ADR-04) — A digest over source version, recipe and encoder build. An encoder upgrade therefore becomes a catalogue-wide cache miss filled lazily by demand, rather than a scheduled re-encode of 250,000 hours.
  5. Bound reclaim loss rather than checkpoint (ADR-05) — Shot-boundary chunks sized so a withdrawn worker costs at most one chunk-minute, with chunk duration as tunable policy. No encoder state is serialised or revived.
  6. Determinism is a release gate (ADR-06) — An encoder build that cannot replay a golden corpus byte-identically is rejected. Determinism is what makes retries free, caches trustworthy and lineage true rather than aspirational.
  7. Lanes with capacity floors, ordered by deadline (ADR-07) — A premiere cannot be starved by a backfill, and a backfill cannot be starved indefinitely. Priority is derived from the deadline and the work remaining, not from a static field that is wrong by the second day.
  8. Batch and on-demand share code, never capacity (ADR-08) — One is throughput-bound on reclaimable compute; the other is latency-bound with reserved headroom. Mixing them would either give a waiting viewer a Spot queue or give the batch fleet an availability requirement it does not need.
  9. The quality gate blocks, and waivers are on the record (ADR-09) — A rendition below the floor is not publishable. The only route past it is a recorded waiver by an authorised approver, and the person under deadline pressure is not the person who sets the floor.
  10. Playback telemetry is a second gate (ADR-10) — A rendition that scores well and plays badly on a real device is quarantined after publication and the manifest revised. Pre-publication metrics are necessary and not sufficient.
  11. Atomic manifest swap is the only publication primitive (ADR-11) — An incomplete ladder publishes as a narrower manifest or not at all. Rollback is a pointer move completing in an assumed 120 seconds with no dependency on the encode fleet.
  12. The ladder is solved per title (ADR-12) — Derived from a complexity profile, with the egress saved weighed against the analysis compute spent and the comparison recorded. At an assumed 40 PB/day of delivered egress, one percent of average bitrate is worth about 400 TB/day.
  13. Residency is an economic decision, re-made continuously (ADR-13) — A rendition is demoted and eventually evicted when its regeneration cost falls below its retention cost over the policy horizon. Eviction is not an end state: a request re-enters the lifecycle at generation.
  14. Recovery is regeneration, not replication (ADR-14) — The standby region carries a warm control plane and no encode fleet. After a region loss the published manifest degrades to the surviving rung set and refills — a visible consequence rather than a hidden one.
  15. The encode fleet is an untrusted zone (ADR-15) — A container escape through a codec parser handling a hostile master is the assumed attack. Workers decode sandboxed with no egress, read one prefix and write one, and content keys never persist in the pipeline.

Why this should still hold up in ten years

Codecs, encoders and cloud price lists will all change inside the life of this design. The decisions above were chosen to be the ones that do not.

  • The long tail is a structural property of catalogues, not a fact about 2026. Viewing concentrates on a small fraction of any large catalogue, and catalogues grow faster than viewing hours per subscriber. The argument for treating renditions as a cache gets stronger as the catalogue grows, not weaker.
  • A new codec is a recipe policy change, not a migration. Because identity is a hash over source, recipe and encoder build, adopting a codec means new recipes and lazy regeneration. The design that pre-encodes everything has to schedule and fund a catalogue sweep for each new codec; this one does not.
  • Reclaimable compute is getting cheaper relative to reserved compute, not dearer. Every provider prices interruptible capacity below committed capacity because it monetises idle fleet. A design whose resilience comes from task granularity rather than worker protection keeps collecting that discount as the gap widens.
  • Determinism and content addressing are what let the system be re-reasoned about later. The expensive failure mode in a ten-year-old media pipeline is not a bad encode; it is a published byte nobody can explain. Recording the recipe and the build digest is cheap now and is the only thing that makes an audit, a regression hunt or a regeneration possible then.
  • The one invariant that must not be traded is protection. Quality, latency and rung availability are all allowed to degrade under failure, and every failure class in this design degrades one of them. Nothing degrades encryption, because that is the trade whose cost arrives as a contract breach rather than a support ticket.

Non-functional targets

Every target below is a stated assumption, invented for a large consumer video-on-demand service so that the architecture has something to be wrong about. The view column points at the diagram where the mechanism that delivers it is drawn.

Quality Target How it is met View
Control-plane availability ≥ 99.95% / month Regional Cloud Run and GKE across three zones; Spanner regional; no dependency on the encode fleet 17
Publish path availability ≥ 99.9% / month Manifest swap is a pointer move in the catalogue projection, independent of packaging and encode 13
Rollback time ≤ 120 s Previous manifest revision retained; rollback is a pointer move with no fleet dependency 19
Encode fleet availability none, by design Spot MIGs; up to 60% withdrawal in 10 minutes absorbed by chunk granularity 14
Catalogue-lane turnaround p95 ≤ 0.80× duration Chunked parallel encode across ~14,000 slots with per-lane capacity floors 14
Premiere launch subset ≤ 45 min for 120 min Three rungs, baseline codec, premiere lane floor; ≥ 160× real-time aggregate 04
Premiere full ladder p95 ≤ 4 h from ingest Remaining rungs added by later manifest revisions 19
On-demand first byte p95 ≤ 1.5 s, p99 ≤ 3.5 s Warm GKE pool, request coalescing, archive range read; nearest-rung fallback past budget 15
Master ingest volume ~126 TB/day 3,500 source hours/day at an assumed 80 Mbps mezzanine; write-once dual-region archive 10
Chunk task throughput 220 k/day, 600 k peak Pub/Sub dispatch with Spanner task claim; no shared mutable state on the encode path 14
Delivered egress ~40 PB/day Per-title ladder; 1% average bitrate ≈ 400 TB/day of egress 12
Top-rung quality ≥ 93 / 100 aggregate Full-reference perceptual score per rung against the source, blocking 18
Worst-window quality no 2 s window < 60 Worst-window score stored with the rendition, not just the aggregate 12
Gate false-pass rate ≤ 0.2% / quarter Post-publication telemetry as a second gate with automatic quarantine 22
Master and recipe durability RPO 0, 7-year lock Dual-region archive buckets with object retention locks; recipes dual-region 11
Control state recovery RPO ≤ 5 s, RTO ≤ 15 min Spanner with a warm standby control plane in the second region 17
Rendition recovery no RPO; refill 30 min p95 Deliberate: loss is a regeneration from master plus recipe 11
Full-ladder encode cost ≤ $1.10 / source hour Open-source encoders on per-second-billed Spot capacity 14
On-demand share of spend ≤ 6% of encode spend Budget check at admission; exceeding it re-tunes eviction rather than raising the cap 15

Scope

In scope

  • Delivery intake by resumable upload, supplier pull or object-store handoff, with integrity verification against a delivery manifest before any work is scheduled.
  • Deep technical probing and conformance classification — accepted, accepted-with-remediation, or rejected — with the applied remediations recorded against the source version.
  • Content complexity analysis and per-title bitrate ladder solving, including codec coverage as a per-title policy and refusal to upscale above the effective source resolution.
  • The recipe plane: immutable, content-addressed, encoder-build-pinned recipes as the only sanctioned encode input, with human overrides recorded with actor and reason.
  • Chunked parallel encode on reclaimable capacity with deterministic, idempotent chunk tasks, lane scheduling with capacity floors, quotas and per-title spend ceilings.
  • Stitching with boundary verification, and segment alignment across every rung of a title so a player may switch rungs at any segment.
  • A blocking quality gate: full-reference perceptual scoring per rung, structural checks, decodability on a device matrix, human sampling, and recorded waivers.
  • Packaging to segmented adaptive formats from one set of elementary streams, common encryption with keys from an external key service, and per-device-class manifests.
  • Atomic manifest publication, launch subsets, rollback, takedown with key revocation, and the demand-tiered residency, eviction and regeneration of every rendition.
  • Lineage from master version through recipe digest and encoder build to published byte, plus the append-only audit trail of overrides, waivers, publishes and rollbacks.

Explicitly out of scope

  • The CDN and edge cache — an adjacent use case in this practice.
  • The player, its adaptive bitrate logic and its device-side decoding.
  • Subtitle and audio-description authoring, translation and timing.
  • Catalogue, artwork and recommendation metadata systems.
  • Live capture, contribution encoding and near-live fast-turnaround paths.
  • The rights and licensing system that decides what may be published where.

What a prototype would have to prove

The decisions in this record are falsifiable, and most of them are falsifiable cheaply. A prototype on a few hundred source hours and a few hundred Spot cores settles the questions that would otherwise be discovered after thirty diagrams have been drawn and a year has been spent. These are the six things it should measure, and the seven scenarios it should survive.

  1. Determinism in practice: that the chosen encoder, wrapped and pinned, produces byte-identical output for the same source, recipe and build across machine types, kernel versions and restarts. If it does not, ADR-04, ADR-05 and ADR-06 all change.
  2. Boundary invisibility: the chunk duration at which stitched output stops showing quality pumping at joins, measured perceptually rather than asserted — this sets the floor for the reclaim-loss bound in ADR-05.
  3. Ladder value: the delivered bitrate saved by per-title solving against the analysis compute spent, on a corpus spanning grainy film, animation and sport, so the trade in ADR-12 has a measured exchange rate.
  4. Regeneration latency: the real archive-restore-plus-encode time for one segment of a top rung, which is the entire basis of the on-demand budget in ADR-08 and ADR-13.
  5. Reclaim behaviour: the actual withdrawal distribution for the chosen Spot machine family in the chosen region, because the chunk-size calculation is only as good as the interval it assumes.
  6. Gate agreement: how often the perceptual floor and a panel of human reviewers disagree, in both directions, which is what tells you whether the floor in ADR-09 is set anywhere near the right place.
  • Sixty percent of the fleet is withdrawn in ten minutes during a premiere encode. The job's progress must stay monotonic, no chunk may be lost twice, and the lane's p95 must hold.
  • A supplier redelivers a master for an already-published title. The published manifest must not change until someone explicitly promotes the new source version.
  • A new encoder build is promoted. No existing rendition may be re-encoded, no existing recipe may change, and the first new recipe must pin the new build.
  • A master is crafted to hang the decoder. The chunk must time out, retries must be bounded, the title must quarantine with a diagnosis, and the lane must keep draining.
  • The key service becomes unavailable mid-packaging. Publication must stall and retry, already-published content must be unaffected, and nothing unencrypted may reach any public path.
  • A long-tail title with every rung evicted becomes popular in one hour. Concurrent requests for the same missing rendition must produce one encode, and nobody may wait past the budget.
  • The encode region is lost. The control plane must fail over inside the RTO, in-flight jobs must restart from completed chunks, and the published manifest must degrade to the surviving rung set rather than break.

Open risks, carried rather than hidden

Risk If it lands Response
The demand prediction for a launch subset is wrong on day zero Premiere viewers get a lower rung in the first minutes — exactly the audience least willing to forgive it, and the one most likely to say so publicly. Keep the launch subset deliberately wider than the prediction suggests for titles above a marketing threshold, and treat the subset's width as a tunable with a cost attached rather than a model output. Question 3 stays open.
The chosen encoder turns out not to be deterministic across machine types Idempotent retries, content addressing and lineage all weaken at once; a duplicate execution stops being free and the cache stops being trustworthy. Make determinism a build-promotion gate before anything depends on it (ADR-06), and pin machine family as well as image digest if the gate cannot otherwise be met.
On-demand generation exceeds its spend ceiling routinely Either the cap is hit and viewers see fallbacks often, or the cap is raised and the economic case for eviction quietly disappears. Treat a sustained breach as a signal that the eviction policy is mis-tuned, not that the cap is wrong; the JIT share of spend is a first-class operational metric for that reason (view 18).
The perceptual gate passes renditions that real devices play badly Quality failures reach viewers, and the first people to notice are the audience rather than the platform. Post-publication telemetry as a second gate with automatic quarantine (ADR-10), and human sampling weighted towards narrow passes and new encoder builds.
A rights holder's contract forbids treating renditions as disposable The critical design decision does not apply to part of the catalogue, and a second, pre-encoded regime appears alongside the first. Make residency a per-title policy from the start so a contractual floor is a policy value rather than an architectural exception, and price the exception explicitly.
Archive-class restore latency makes the on-demand budget unreachable The cache model's cost advantage survives but its viewer-facing promise does not, and the fallback becomes the normal path rather than the exception. Measure restore latency in the prototype before committing; keep a hot copy of the master for titles above a demand threshold, which is a cheaper exception than keeping every rendition.

The reasoning behind every component and technology choice is in the Architecture Decision Record: 15 records across 5 areas, each with the alternatives that lost and what the choice costs.

Document 2 of 2 · 63 min read

Architecture Decision Record

Streaming Video Encoding & Packaging Pipeline · Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records

The argument these decisions serve is summarised in the Architecture One-Pager.

Fifteen decisions, grouped into five areas. Each carries the question that forced it, how it is realised on Google Cloud, what was rejected, what would flip it, and why it should still hold in ten years.

Status of this document. Evidence note. Every quantity in this package is a stated assumption for a large consumer video-on-demand service, invented so the architecture commits to something arguable. None is taken from a published figure for any named streaming service, and no claim is made that any of them reflects a real operator's numbers. Where a number drives a decision, the decision says so and names the number. Replace each one with a measured value before building anything from this.

How to read a record

  • Question: The forcing question: why a decision was needed at all.
  • Context: The requirement, the scale and the constraint that make it hard.
  • Decision: What this architecture does, stated so it can be checked.
  • How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
  • Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
  • Consequences: What the choice buys and what it costs, both kept visible.
  • Choose differently when: The conditions that would flip the decision for your system.
  • Why it holds up over time: What keeps the decision right as scale, staff and technology change.
  • Lesson: The principle that transfers beyond this platform.

Decision map

custody-and-identity: What the system actually owns, and how a derived artefact is named.

  • ADR-01 · The master plus the recipe is the asset; every rendition is a cache entry
  • ADR-02 · The mezzanine is immutable on arrival, and a redelivery is a new source version
  • ADR-03 · The recipe is the only sanctioned encode input, pinned to an encoder build
  • ADR-04 · Renditions are identified by content hash, not by assigned identity

encode-execution: Turning a recipe into bytes on compute that is being taken away.

  • ADR-05 · Bound reclaim loss with small chunks rather than checkpointing encoder state
  • ADR-06 · Determinism is an encoder release gate, not a hoped-for property
  • ADR-07 · Priority lanes with capacity floors, ordered by deadline and work remaining
  • ADR-08 · Batch and on-demand generation share code, never capacity

quality-and-publication: Proving an output is good enough, and putting it in front of people.

  • ADR-09 · The quality gate blocks publication, and a waiver is a recorded act
  • ADR-10 · Playback telemetry is a second, post-publication gate
  • ADR-11 · Atomic manifest swap is the only publication primitive

economics-and-lifecycle: Deciding which renditions are worth existing, and for how long.

  • ADR-12 · The bitrate ladder is solved per title from a complexity profile
  • ADR-13 · Rendition residency is an economic decision, re-made continuously

protection-and-operations: Keeping an unreleased master and a content key safe, and recovering a region.

  • ADR-14 · Regional recovery is regeneration from custody, not rendition replication
  • ADR-15 · The encode fleet is an untrusted zone, and no fallback weakens protection

Technology by capability

Google Cloud carries 7 of the 56 existing documents in this practice against Amazon Web Services' 10, open-source-on-premises' 11 and Microsoft Azure's 8, so it is the honest rotation — and the topic suits it. The central question here is cost per delivered hour, and that only becomes a real design question when the encoder is yours and the compute is reclaimable: per-second-billed preemptible instances, a managed batch scheduler, object storage with lifecycle classes, and a media-aware CDN downstream. A managed per-minute transcoding API would answer the interesting question with a price list. The requirement document stays vendor-neutral throughout; the choices below belong to this architecture, not to the ask, and each row names what it was chosen over.

Capability Choice Origin Credible alternative Why this one Record
Master custody — immutable, retention-locked, dual-region Cloud Storage dual-region bucket, Archive class, object retention lock Google Cloud Single-region with manual replication; a self-managed object store on-premises Dual-region gives RPO 0 without a replication job to operate, and retention lock is a storage-level guarantee rather than application logic the pipeline could be talked out of. ADR-02
Recipe store — immutable, content-addressed, indexed Cloud Storage dual-region for the documents, Spanner for the index Google Cloud A relational store holding the documents inline; a document database The documents are immutable and want object-store durability; only the lookup wants a transactional index. Splitting them keeps the custody artefact out of a schema that will change. ADR-03
Job and chunk-task state — transactional, cross-zone Cloud Spanner, regional Google Cloud PostgreSQL with a leader per region; a lease protocol over a key-value store This is the one place in the design needing transactional correctness across zones — task claim, lease expiry and lane accounting in one transaction. The alternative is a lease protocol we would have to write and prove. ADR-07
Batch encode fleet — reclaimable, per-second billed Compute Engine Spot VMs in zonal managed instance groups, gVisor around the decoder Google Cloud Committed-use reserved instances; a managed transcoding API priced per minute Reclaimable capacity is the economic premise of the whole design, and per-second billing is what makes small chunks affordable. A managed per-minute API removes the decision this use case exists to make. ADR-05
Encoders and packaging Open-source encoders (FFmpeg-family, SVT-AV1) and Shaka Packager on GKE Open source A managed transcoding and packaging service; a commercial encoder licence Owning the encoder is what makes determinism gateable, build pinning enforceable and the ladder solvable per title. Common encryption from one elementary-stream set is a packager requirement, not a service feature. ADR-06
On-demand generation — latency-bound, reserved GKE regional warm node pool beside a Cloud Run packaging origin Google Cloud Scale-to-zero serverless; preemption inside the batch fleet; edge compute An assumed p95 of 1.5 s to first byte rules out a cold start and rules out queuing behind Spot work. Edge compute is left open in Question 7 rather than chosen. ADR-08
Task dispatch — at-least-once, high fan-out Pub/Sub, with task claim in Spanner Google Cloud A broker with exactly-once semantics; direct scheduler-to-worker assignment Determinism and idempotent chunks make at-least-once sufficient, so the expensive delivery guarantee buys nothing. The transactional claim lives where the state already is. ADR-06
Content keys and encryption at rest Cloud KMS for CMEK and content-key generation, keys held in memory for one packaging job Google Cloud A self-hosted key manager; keys cached in the packaging tier for throughput The pipeline must never be a place a content key is stored, and per-rights-holder key separation is a contract requirement for part of the catalogue. Caching keys would trade a contractual exposure for a latency gain nobody asked for. ADR-15
Quality, cost and playback analytics BigQuery, partitioned by publish date, fed from Pub/Sub Google Cloud A time-series database; a warehouse in the analytics estate Gate decisions, metric scores and playback outcomes are append-only, queried by date and joined against digests — a columnar store is the natural shape, and the post-hoc gate needs to query it directly. ADR-10
Rendition residency and eviction Cloud Storage regional with lifecycle rules, driven by a demand-tier controller on GKE This design Pure age-based lifecycle rules; manual tiering; no eviction at all The eviction rule is economic rather than temporal — regeneration cost against retention cost over the policy horizon — and an age rule cannot express it. Lifecycle rules execute the decision; the controller makes it. ADR-13

The decisions, and the alternatives that lost

custody-and-identity

What the system actually owns, and how a derived artefact is named.

ADR-01 · The master plus the recipe is the asset; every rendition is a cache entry

Status: Accepted · Shown on views: 11, 15, 19

Is a published rendition a durable artefact the platform is obliged to preserve, or a derived file it may delete and rebuild whenever that is cheaper?

Context. A large catalogue produces a brutal asymmetry. An assumed 400,000 titles and 250,000 source hours, each wanting a dozen renditions across more than one codec, is on the order of 1.2 million rendition sets. An assumed 72% of catalogue hours attract under 0.1% of total watch time. Treating every rendition as a publication means paying regional storage, lifecycle management and — if they are treated as irreplaceable — cross-region replication and backup, in perpetuity, for files that will be requested a handful of times or never. The alternative is uncomfortable in a different way: if renditions can be deleted, a viewer can ask for one that does not exist, and somebody has to answer them inside a latency budget. Nothing else in the design can be settled before this is, because it determines what gets replicated, what gets an RPO, what disaster recovery means, what a codec upgrade costs, and whether there is an on-demand encoding fleet at all.

Decision. The mezzanine master and the recipe are the only assets. Renditions and segments carry no RPO, no backup and no cross-region replication; they are evictable under an economic policy and regenerable from the master and the recipe. On-demand regeneration is a first-class, latency-budgeted, separately scaled component of the architecture rather than a repair tool.

How it is realised on AWS. Masters land write-once in a dual-region archive bucket with object retention locks, and recipes sit alongside them in a dual-region bucket indexed in Spanner. Renditions and segments live in a single-region bucket, content-addressed, with lifecycle rules that demote and delete by demand tier. The storage-zones view draws exactly three zones — irreplaceable, recoverable, disposable — and the disposable zone's RPO is written on the diagram as "none, by design" so that nobody later adds replication to it believing they are fixing an oversight. A warm GKE pool beside the packaging origin serves regeneration with request coalescing on the rendition digest and a nearest-resident-rung fallback.

Option Verdict Reasoning
Renditions are a cache over master plus recipe Chosen Storage becomes a per-title economic decision; DR becomes replication of two small things plus a regeneration budget; the cost is a viewer-visible fallback on cold titles.
Pre-encode the full ladder and preserve every rendition Right elsewhere Right for a small catalogue, a premium archive, or an operator whose contracts forbid disposing of deliverables. Uniform latency, no fallback to explain, and a storage bill that scales with catalogue rather than viewing.
Pre-encode everything but back up nothing Rejected Keeps the full storage bill and still has to regenerate after a loss — the worst of both, with no on-demand path built because nobody planned for one.
Encode only on first request, with nothing pre-encoded Rejected Makes the premiere the failure case. The first hour of a new release is precisely when nothing may be generated on demand.

What it buys

  • Disaster recovery for 1.2 million rendition sets reduces to replicating masters and recipes — orders of magnitude smaller — plus a measured regeneration budget.
  • A bad encode is a recipe revision rather than a data-repair exercise, and a codec upgrade is a policy change applied lazily as demand arrives.
  • The long tail stops being a perpetual storage obligation and becomes a retention-versus-regeneration calculation that can be re-made as prices move.
  • Storage class, eviction threshold and launch-subset width all become tunable policy values with a visible cost, rather than architectural facts.

What it costs

  • A real viewer can experience a cache miss, and the honest p99 of that miss is an assumed 3.5 seconds to first byte.
  • A nearest-rung fallback is visible: somebody will occasionally watch a lower rung than their connection deserves, and that has to be acceptable to the product.
  • The architecture now needs a second encoding fleet with its own latency SLO, its own reserved headroom and its own spend ceiling.
  • Any rights holder contract that forbids disposing of deliverables becomes an exception that must be priced and policed per title.

Choose differently when. Three things would flip it. If storage became cheap enough relative to compute that retaining a full ladder for the whole catalogue cost less than the on-demand fleet plus its fallbacks, pre-encoding wins outright. If archive-class restore latency proves unreachable within the first-byte budget even for a warm pool, the cache model keeps its cost advantage but loses its viewer-facing promise, and the honest response is a hot master copy for high-demand titles rather than retaining renditions. And if a majority of the catalogue arrived under contracts that forbid treating deliverables as disposable, the exception would be the rule and the decision should be reversed rather than special-cased.

Why it holds up over time. The decision rests on a property of catalogues rather than a property of 2026: viewing concentrates, and catalogues grow faster than viewing hours per subscriber. As a catalogue grows the case strengthens. It also gets stronger as codecs multiply, because every new codec multiplies the rendition count under the pre-encode model while costing only new recipes under this one. The main thing that would date it is a step change in storage economics, and storage has historically fallen in price more slowly than codec counts have risen.

Lesson. Decide what is genuinely irreplaceable before deciding anything about durability. Most of what a pipeline produces is a function of something smaller, and naming that function is cheaper than preserving its output.

ADR-02 · The mezzanine is immutable on arrival, and a redelivery is a new source version

Status: Accepted · Shown on views: 10, 12, 06

When a supplier sends a corrected master for a title that is already published, does it replace the old one, or become a new thing that someone has to promote?

Context. Redelivery is routine: a wrong audio mix, a missing frame, a mastering error, a different territorial cut. The tempting behaviour is to overwrite, because the supplier thinks of it as a correction and the file has the same name. Overwriting breaks three things at once. It invalidates every rendition derived from the old bytes without any record that it did so; it makes the lineage claim false, because the published byte can no longer be traced to the master that produced it; and it means a supplier can change what viewers see without anyone at the platform deciding that they should.

Decision. Received bytes are immutable on arrival: never modified in place, never overwritten by a later delivery. Integrity is verified against a delivery manifest before any work is scheduled. A redelivery becomes a new source version, and replacing what is published requires an explicit promotion.

How it is realised on AWS. The upload endpoint writes to a write-once prefix in the dual-region archive bucket under a key derived from the content checksum, with object retention locks preventing modification for the retention term. The source_version row carries the checksum as a unique key, the full technical profile from the deep probe, the conformance verdict and the applied remediations. Promotion is a separate, audited operation that creates new recipes; it does not mutate the existing ones. Conformance runs before scheduling so a rejection reaches the supplier within an assumed one hour rather than after a day of encoding.

Option Verdict Reasoning
Immutable landing, redelivery as a new version, explicit promotion Chosen Lineage survives, the platform decides what is published, and a bad redelivery cannot silently invalidate a working title.
Overwrite in place and re-encode Rejected Cheap to implement and destroys traceability. Also makes rollback impossible, because the bytes the current renditions came from no longer exist.
Versioned bucket with the latest version implicitly authoritative Rejected Keeps the history but still lets a supplier's upload change what is published. The storage layer should not hold an editorial decision.
Accept any delivery and let the quality gate catch problems Rejected Spends a full ladder of encode on a master that a checksum would have rejected in seconds, and tells the supplier far too late to redeliver before air.

What it buys

  • Every published byte traces to a specific master version, which is what makes a regression hunt or an audit possible years later.
  • Rollback is real: the previous source version and its recipes still exist, so reverting is a promotion rather than a recovery.
  • A redelivery cannot invalidate a live title as a side effect; somebody has to choose, and the choice is audited.
  • Rejecting on integrity before scheduling saves an assumed full-ladder encode on every malformed delivery and gets the supplier a verdict inside an hour.

What it costs

  • Storage holds every version of every master for the retention term, including the ones that were superseded the same week.
  • Operations carries a promotion step that someone has to perform, and a queue of unpromoted redeliveries that someone has to watch.
  • Suppliers who expect overwrite semantics need an explanation, and some will need a contract amendment.
  • The checksum-derived key means a byte-identical redelivery is a no-op, which is correct but occasionally confusing to a supplier who expects to see a new record.

Choose differently when. If retention cost for superseded masters became material — a catalogue dominated by high-bitrate redeliveries, say — a tiering rule that deletes unpromoted source versions after a window would be a reasonable amendment, keeping immutability while bounding the history. The decision would only genuinely flip if the platform stopped being the publisher of record, for instance if a supplier operated the publication decision themselves, at which point immutability on arrival is still right but promotion moves to them.

Why it holds up over time. Immutability on write has become the default expectation for object storage and for anything auditable, and retention locks are now a standard storage feature rather than an add-on. The decision gets easier to implement over time, not harder. The underlying reason — that an editorial decision should not be expressible as a file overwrite — does not depend on any technology.

Lesson. Make the system of record immutable, and make the act of changing what people see a separate, named, audited operation. Those two things are not the same and should not share a mechanism.

ADR-03 · The recipe is the only sanctioned encode input, pinned to an encoder build

Status: Accepted · Shown on views: 07, 12, 16

Where do the encoding parameters for a title live — in the job, in the worker's configuration, or in a separate immutable artefact the job merely references?

Context. Parameters have a way of accumulating in whichever place is most convenient to change: a job submission field, a per-worker config file, an environment variable set during an incident. Each of those makes the parameters un-reconstructable afterwards. Meanwhile the one derived artefact in the whole pipeline that cannot be recomputed from anything else is the decision about how a title should be encoded: the analysis that informed it, the ladder that came out, the codec policy, and the encoder build it was meant for. Everything else — chunks, renditions, segments, manifests — is a function of the master and that decision.

Decision. The recipe is an immutable, versioned, content-addressed document holding the ladder, the analysis inputs, every encoder parameter and the encoder build identifier. It is the only sanctioned input to an encode. No out-of-band parameter may reach a worker, and a human override is recorded with actor and reason or refused.

How it is realised on AWS. Recipes are JSON documents keyed by a digest over source version, analysis output, encoder build and parameters, stored in a dual-region bucket with a Spanner index. Chunk tasks carry the digest, not the parameters. A worker fetches the recipe by digest and refuses the task if its own encoder build does not match the pinned one. The layered-architecture view puts the recipe plane above the control plane and shades it, because its lifetime and its custody requirements are those of the master rather than those of a job.

Option Verdict Reasoning
Immutable content-addressed recipe, build-pinned, as the only input Chosen Makes an encode reproducible years later and makes lineage a lookup rather than an archaeology exercise.
Parameters carried on the job record Rejected Job records get pruned, and parameters that live inside a mutable row stop being able to explain a published byte once the row changes.
Parameters in worker configuration, deployed with the fleet Rejected Makes an encode a function of when it ran. Two chunks of the same rendition can then differ, which is the defect that is hardest to find.
A ladder template per content category, resolved at encode time Deferred A reasonable optimisation for the solver's inputs, but the resolved result must still be frozen into a recipe before any chunk runs.

What it buys

  • An encode is reproducible from the recipe digest alone, which is what makes retries free and regeneration possible eighteen months later.
  • Lineage from published byte to master version is a two-hop lookup rather than an investigation.
  • An incident cannot be 'fixed' by changing a parameter on a worker, because the worker has no parameters to change.
  • Overrides become visible: the audit log shows who widened a ladder and why, which is the only way to learn whether overrides are a symptom.

What it costs

  • Every parameter change produces a new recipe version, so the recipe store grows with experimentation as well as with the catalogue.
  • The solver has to be complete: anything it forgets to record is a parameter that cannot be reconstructed, and the gap will not be obvious.
  • Operators lose the ability to make a quick fleet-wide tweak, which is occasionally genuinely inconvenient during an incident.
  • A build-pinned recipe means the fleet must be able to run older builds, so build images have to be retained as long as the recipes that reference them.

Choose differently when. If encoder builds could not be retained for the life of their recipes — a licensing restriction, or an upstream project that deletes releases — build pinning becomes unenforceable and the honest fallback is to record the build identifier descriptively and accept that regeneration may be equivalent rather than identical. That weakens ADR-04 and ADR-06 with it, so it should be resisted rather than accommodated quietly.

Why it holds up over time. Content-addressed immutable configuration is the direction every build, deployment and data system has moved for two decades, for the same reason: it is the only way to answer what produced a given output. The specific parameters will change with every codec generation; the decision to freeze them in an addressable artefact will not.

Lesson. Find the one derived artefact in your pipeline that cannot be recomputed, give it custody and an address, and make everything else a function of it.

ADR-04 · Renditions are identified by content hash, not by assigned identity

Status: Accepted · Shown on views: 12, 16, 19

Is a rendition named by a digest of what produced it, or by an identifier the platform assigns to the slot it fills?

Context. This looks like a naming convention and is actually the decision that prices every future codec and encoder migration. Under a content hash over source version, recipe and encoder build, promoting a new encoder build changes the digest of everything it would produce, which means the whole catalogue becomes a cache miss. Under an assigned identity — title, rung, codec — the upgrade is invisible: the same slot is simply filled by newer bytes next time, and the question "what produced this byte" becomes an inference from timestamps. The first is correct and potentially enormous. The second is cheap and quietly lies.

Decision. A rendition is identified by a digest over (source version, recipe digest, encoder build). An encoder upgrade therefore makes the catalogue a cache miss, which is affordable only because of ADR-01: the miss is filled lazily, by demand, and never swept.

How it is realised on AWS. rendition_digest is the primary key of the rendition table and the object key prefix in the rendition bucket. Segments are content-addressed under their rendition. A manifest references digests, so a player's request names exactly what the recipe would produce. The build-rollout view makes the consequence explicit: a promoted build is pinned by new recipes only, existing recipes and renditions are untouched, and the catalogue migrates as demand asks for renditions that do not exist yet.

Option Verdict Reasoning
Content hash over source, recipe and encoder build Chosen Correct, traceable, and affordable only in combination with lazy regeneration. Makes an upgrade a demand-driven migration with no sweep to schedule or fund.
Assigned identity per (title, rung, codec) Right elsewhere Right where renditions are pre-encoded and preserved: the upgrade is a no-op and the slot semantics match the storage model. Costs the ability to say what produced a byte.
Content hash excluding the encoder build Rejected Makes the upgrade free and makes two bytes with the same digest potentially different, which destroys the property the digest exists for.
Hybrid: assigned identity with the build recorded as metadata Rejected Traceability becomes advisory. Once a digest is not derived from the build, nothing enforces that the metadata is right.

What it buys

  • A codec or encoder migration costs nothing to schedule: there is no sweep, no migration plan and no catalogue-wide re-encode budget.
  • Mixed-build output within a title is impossible by construction, because every rendition's digest names its build.
  • A cache is trustworthy: if the digest matches, the bytes are the bytes the recipe specifies, so no verification step is needed on a hit.
  • Regression hunting is a query — which builds produced the quarantined renditions — rather than a bisection.

What it costs

  • Promoting a build invalidates nothing and misses everything: every subsequent regeneration is a new encode, so build promotions have a diffuse compute cost that is hard to attribute.
  • Digests are opaque, so every human-facing surface needs a translation layer to say which title and rung a digest belongs to.
  • A title can hold renditions from several builds simultaneously during a long migration, which is correct and complicates quality comparisons.
  • Build images must be retained for as long as any recipe references them.

Choose differently when. If encoder builds were promoted very frequently — weekly, say — the diffuse regeneration cost could overwhelm the saving, and the right answer becomes batching promotions to a slower cadence rather than changing the identity scheme. If the design ever abandoned ADR-01 and pre-encoded everything, assigned identity becomes the better choice immediately, because a catalogue-wide cache miss with no lazy fill is a catalogue-wide re-encode.

Why it holds up over time. Content addressing has won in every adjacent domain — container images, build artefacts, package registries, version control — for the same reason it wins here. Codec generations will keep arriving, roughly one significant one every few years, and each arrival makes a lazily-migrating identity scheme more valuable than a scheme that requires a funded sweep.

Lesson. A naming scheme is a migration policy in disguise. Decide what you want an upgrade to cost, then pick the identity that produces that cost.

encode-execution

Turning a recipe into bytes on compute that is being taken away.

ADR-05 · Bound reclaim loss with small chunks rather than checkpointing encoder state

Status: Accepted · Shown on views: 14, 13, 22

When the platform loses a worker mid-encode, does it lose one small unit of work, or does it preserve the work by serialising and reviving the encoder's state elsewhere?

Context. The economic premise of the whole design is reclaimable capacity: an assumed 18% hourly reclaim at steady state and up to 60% of the fleet withdrawn inside a ten-minute window. At that rate a whole-file encode of a feature film would rarely finish. Two mechanisms can rescue it. Splitting the source into independently encodable chunks bounds the loss to one chunk but multiplies per-task setup cost and, more importantly, multiplies the number of boundaries at which rate control restarts — which is the cause of visible quality pumping at joins. Checkpointing preserves the work but requires the encoder to serialise its internal state and another machine to revive it, which most production encoders cannot do and none do portably.

Decision. Split on shot or GOP boundaries into chunks sized so expected reclaim loss stays within a stated bound — an assumed one chunk-minute — and reschedule a lost chunk without operator action. Chunk duration is a tunable policy value, not a constant. No encoder state is serialised or revived.

How it is realised on AWS. The recipe carries the chunk plan derived from shot boundaries, so chunk edges land where a cut already breaks temporal prediction and the boundary is cheapest to make invisible. Chunk tasks are rows in Spanner claimed under a lease; lease expiry is the reclaim signal, and a requeue is one row update. Partial output in chunk scratch is discarded rather than salvaged. The swimlane view draws Reclaim as one of six columns precisely because it is a normal stage of the flow rather than an error path.

Option Verdict Reasoning
Shot-boundary chunks with bounded loss and reschedule Chosen Works with any encoder, needs no state serialisation, and puts the quality cost where a cut already is. Chunk size becomes the dial between loss and boundary count.
Checkpoint encoder state and resume on another worker Deferred Preserves work and is the right answer for very long or very complex sources where chunking costs too much quality. Phase 3, and only if the chosen encoder can serialise state at all.
Whole-file encode on reserved capacity Right elsewhere Right for a small volume of premium titles where boundary artefacts are unacceptable and the compute premium is affordable. Does not scale to 3,500 source hours a day.
Fixed-duration chunks irrespective of content Rejected Simplest to schedule and puts boundaries in the middle of motion, which is exactly where a restart of rate control is most visible.

What it buys

  • Job progress is monotonic across any number of reclaims, so a premiere encode survives a withdrawal storm without operator involvement.
  • The mechanism is encoder-agnostic, so the design is not hostage to one encoder's feature set.
  • Chunk duration gives a single dial that trades reclaim loss against boundary count, measurable rather than argued.
  • Parallelism across thousands of chunks is what makes the assumed 45-minute premiere turnaround arithmetically possible at all.

What it costs

  • Every boundary is a potential artefact, and verifying boundaries is a mandatory structural check rather than an optional one.
  • Per-chunk setup — container start, recipe fetch, source range read — is paid thousands of times per title and is a real fraction of the bill for short chunks.
  • Rate-control continuity across chunks has to be handled by the encoding strategy, which constrains which encoder settings are usable.
  • Chunk scratch holds partial outputs with a TTL, which is storage spent on work that will sometimes be thrown away.

Choose differently when. Two measurements would flip it. If the prototype finds no chunk duration at which boundaries are perceptually invisible for grainy or high-motion content, checkpointing moves from Phase 3 to mandatory for that content class. And if reclaim intervals turned out to be much longer than assumed — a region and machine family with 2% hourly reclaim rather than 18% — larger chunks or whole-file encodes become viable and the setup overhead argument reverses.

Why it holds up over time. The decision depends on reclaimable compute being materially cheaper than committed compute, which is structural: providers price interruptible capacity to monetise idle fleet, and that gap has widened rather than narrowed. Encoders have also not converged on portable state serialisation in twenty years, so the alternative has not become easier. What will change is the chunk size, as machine types and codecs change — which is why it is a policy value.

Lesson. When the platform cannot stop losing workers, buy resilience with task granularity rather than with worker protection — and then treat the granularity as a measured dial rather than a constant someone once chose.

ADR-06 · Determinism is an encoder release gate, not a hoped-for property

Status: Accepted · Shown on views: 16, 14, 12

Does the platform require byte-identical output for the same source, recipe and build, and does it refuse to promote an encoder that cannot deliver it?

Context. Three separate properties of this architecture quietly assume determinism. Idempotent retries assume a duplicate execution is indistinguishable from a single one, which is only true if the output is identical. Content-addressed identity assumes a digest over inputs predicts the output, which is only true if the mapping is a function. And lineage assumes that recording a recipe and a build explains a published byte, which is only true if those inputs fully determine it. Encoders are full of things that break this: thread-count-dependent partitioning, time-based decisions, hardware-specific instruction paths, non-deterministic rate-control heuristics. None of them announce themselves, and the symptom — two chunks of one rendition that differ subtly — is among the hardest defects to localise.

Decision. An encoder build is promoted only if it replays a golden corpus byte-identically across the machine types the fleet runs. A build that cannot is rejected before anything else about it is measured. Quality regression against the current build is a separate, subsequent gate.

How it is realised on AWS. The rollout flow puts the determinism gate immediately after the build, before the regression gate, the shadow encode and the canary. The gate re-encodes a fixed corpus spanning grain, animation, high motion and HDR, on every machine family in the fleet, and compares digests. The build registry holds only approved digests, and a worker whose running build does not match the recipe's pinned build refuses the task rather than producing output nobody can reproduce.

Option Verdict Reasoning
Determinism as a hard promotion gate Chosen Protects idempotence, content addressing and lineage in one step, and catches the failure in CI rather than in a year-old rendition.
Best-effort determinism with equivalence checking at use Rejected Means verifying every cache hit, which removes the cache's entire benefit, and turns every regeneration into a quality comparison.
Pin machine family as well as build to obtain determinism Deferred A legitimate fallback if a strongly preferred encoder is deterministic per machine family but not across them. Costs scheduling flexibility, which costs Spot availability.
Accept non-determinism and use assigned identity instead Right elsewhere Coherent for a pipeline that pre-encodes and preserves renditions, where nothing depends on reproducing a byte. Not coherent with ADR-01 or ADR-04.

What it buys

  • A duplicate chunk execution is free, so at-least-once dispatch is sufficient and no distributed lock is needed on the hot path.
  • A cache hit needs no verification: a matching digest means the bytes are the specified bytes.
  • Lineage claims are true rather than aspirational, which is what makes an audit or a regression hunt possible years later.
  • The failure is caught in a build pipeline against a fixed corpus, where it is cheap, instead of in a stitch defect report.

What it costs

  • Some otherwise excellent encoders and some hardware-accelerated paths will fail the gate and be unavailable.
  • Encoder settings that improve quality through non-deterministic heuristics are off the table, which may cost measurable bitrate.
  • The golden corpus and the cross-machine-type replay are real infrastructure to build and maintain.
  • Upstream encoder releases may be blocked for months while determinism is fixed, so the platform tracks upstream more slowly.

Choose differently when. If the best available encoder for a mandatory codec were irreducibly non-deterministic, the gate becomes unenforceable for that codec. The ordered fallbacks are: pin machine family; then accept per-codec equivalence checking with a documented loss of cache trust; and only then reconsider ADR-04. Reversing the gate wholesale without reversing ADR-01 and ADR-04 would leave three decisions resting on an assumption that is no longer true.

Why it holds up over time. Reproducible builds have moved from a research interest to a baseline expectation across packaging, container and supply-chain tooling in a decade, and the same pressure is arriving in media tooling. The requirement is more likely to become easier to satisfy than harder. What will not change is that three of this architecture's load-bearing properties are consequences of it.

Lesson. If several decisions quietly depend on one property, make that property a gate rather than an assumption — and put the gate where failing it is cheap.

ADR-07 · Priority lanes with capacity floors, ordered by deadline and work remaining

Status: Accepted · Shown on views: 14, 13, 06

How does a day-and-date premiere get through a fleet that is simultaneously grinding through a catalogue backfill, without the backfill being starved forever?

Context. The workload has two shapes that meet in one fleet. A premiere is small, urgent and has an immovable embargo. A backfill is enormous, patient, and will consume every available slot if allowed to. Pure priority ordering starves the backfill indefinitely, which sounds acceptable until a rights window expires on an unencoded title. Pure fairness misses premieres. Static priority fields drift: a title marked urgent six weeks ago is not urgent today, and a title nobody flagged is on air tomorrow. Meanwhile the fleet's own capacity is varying continuously because it is reclaimable, so any scheduling decision has to survive the fleet halving under it.

Decision. Work is admitted through named lanes with guaranteed minimum capacity floors, so a premiere lane cannot be starved by a backfill lane and a backfill lane cannot be starved indefinitely. Within a lane, ordering is derived from the per-title deadline and the work remaining rather than from a static priority field. Per-owner quotas and per-title spend ceilings are enforced at admission as hard stops.

How it is realised on AWS. Lane accounting lives in Spanner beside the chunk tasks, so a claim and a lane decrement are one transaction. The scheduler runs on regional GKE and re-derives ordering continuously from deadline minus estimated remaining chunk-seconds, which means a job that is falling behind rises automatically and a job with slack falls. Floors are expressed as a share of currently available slots rather than an absolute count, so a reclaim storm shrinks every lane proportionally instead of collapsing the smallest. Progress is published as completed chunk-seconds against total so an operator sees a slip coming while there is still time to act.

Option Verdict Reasoning
Named lanes with capacity floors, deadline-derived ordering within a lane Chosen Bounds both failure modes, and makes urgency a computed property that cannot go stale.
Strict priority ordering Rejected Simple and starves the backfill until a rights window expires on a title nobody was watching the queue for.
Fair-share scheduling across owners Rejected Correct for multi-tenant fairness and wrong for deadlines: it has no concept of an embargo at midnight.
Separate fleets per lane Right elsewhere Right where lanes have genuinely different hardware needs or isolation requirements. Here it fragments Spot capacity, which is the scarce resource.

What it buys

  • A premiere and a catalogue migration coexist in one fleet with a bounded worst case for each.
  • Urgency cannot go stale, because it is derived from the deadline rather than asserted once at submission.
  • Floors expressed as a share of available slots mean a reclaim storm degrades every lane proportionally rather than destroying one.
  • Spend ceilings at admission stop a runaway recipe before the money is spent rather than alerting afterwards.

What it costs

  • The scheduler needs a remaining-work estimate per job, and a bad estimate mis-orders the queue in a way that is hard to notice.
  • Lane floors are a tuning surface that will be argued about, and a wrong floor is invisible until the month a premiere is late.
  • Deriving order continuously means the queue is not stable, which makes 'why did my job not run' a harder question to answer.
  • A hard spend stop will occasionally block legitimate work at an inconvenient moment, and that is the intended behaviour.

Choose differently when. If the fleet were large enough relative to the workload that contention disappeared, lanes become bookkeeping and strict priority would do. If the opposite — a persistently over-subscribed fleet — the floors stop being floors in any meaningful sense and the real answer is capacity, not scheduling. And if deadlines proved unreliable as data, because nobody maintains the release calendar, derived ordering is no better than a static field and lanes would carry the whole weight.

Why it holds up over time. Deadline-derived scheduling over interruptible capacity is the stable shape for any workload mixing urgent small jobs with patient large ones, and it long predates this design. The specific floors and the estimator will be re-tuned continuously; the decision to compute urgency rather than declare it is what survives.

Lesson. Priority that is asserted once is wrong by the second week. Derive urgency from the thing that is actually true — a deadline and the work left — and bound the starvation you are willing to accept in both directions.

ADR-08 · Batch and on-demand generation share code, never capacity

Status: Accepted · Shown on views: 15, 17, 08

Does a viewer waiting for a rendition that does not exist get served by the same fleet that is grinding through the backfill, or by a separate one?

Context. ADR-01 creates a second encoding workload with completely different properties. Batch encode is throughput-bound, tolerant of minutes of queuing, and economically dependent on reclaimable capacity. On-demand generation is latency-bound with an assumed p95 of 1.5 seconds to first byte, and a viewer is sitting in front of it. Running both on one fleet means one of two bad outcomes: either the waiting viewer joins a Spot queue behind 200,000 backfill chunks, or the batch fleet inherits an availability and latency requirement that destroys the Spot economics that justified it.

Decision. The two paths share the encoder, the recipe and the code, and share no capacity. Batch runs on reclaimable Spot managed instance groups with no availability target. On-demand runs on a warm, reserved GKE pool beside the packaging origin, with request coalescing on the rendition digest, a spend ceiling of an assumed 6% of total encode spend, and a nearest-resident-rung fallback when the budget cannot be met.

How it is realised on AWS. The deployment view separates them physically: zonal Spot MIGs for batch, a regional warm pool for on-demand, with the origin and the manifest cache in the same latency-bound tier. Both consume the same recipe by digest, which is what makes the shared-code claim true rather than nominal — an on-demand regeneration produces the same digest the batch fleet would have produced. Coalescing happens at the origin, so a hundred concurrent requests for a missing rendition produce one encode. The budget check sits at admission to the warm pool, and exceeding it serves a fallback rather than queuing.

Option Verdict Reasoning
Shared code, separate capacity, separate SLOs Chosen Each workload gets the capacity model its requirement implies, and determinism guarantees the outputs are interchangeable.
One fleet with priority preemption for on-demand work Rejected Preempting a Spot worker that may itself be reclaimed gives the viewer a queue with two sources of delay and no bound.
On-demand generation at the edge Deferred Attractive for first-byte latency and currently incompatible with master access and key handling. Question 7 keeps it open.
No on-demand path; pre-encode everything Right elsewhere Right under a pre-encode-and-preserve model. Reverses ADR-01, and with it the storage and DR economics.

What it buys

  • A waiting viewer never queues behind batch work, and the batch fleet keeps its Spot pricing and its absence of an availability target.
  • Determinism makes the two paths' outputs identical, so a rendition's provenance does not depend on which fleet produced it.
  • The on-demand path has its own spend ceiling, which turns an eviction-policy mistake into a measurable signal instead of a surprise invoice.
  • Coalescing bounds the cost of a sudden spike on a cold title to a single encode.

What it costs

  • Reserved warm capacity is paid for whether or not it is used, and sizing it is a forecasting problem with no good data on day one.
  • Two fleets mean two sets of operational behaviour, two scaling policies and two things to be on call for.
  • The fallback contract — serve a lower rung rather than wait — has to be acceptable to the product, and that is a conversation, not a configuration.
  • Archive-class master restore sits inside the first-byte budget, which couples a latency SLO to a storage class decision.

Choose differently when. If on-demand demand were low and bursty enough that a serverless cold start fit inside the budget, the reserved pool becomes waste and the right answer is scale-to-zero. If it were high and steady, the pool stops being an exception and the honest conclusion is that eviction is too aggressive — which is why the JIT share of spend is an operational metric rather than a finance line.

Why it holds up over time. Separating latency-bound from throughput-bound work onto different capacity is one of the most stable patterns in systems design, and nothing about media encoding makes it less true. What will move is the boundary: as cold-start times fall and edge compute matures, the warm pool may become serverless or move outward, and both are changes of realisation rather than of decision.

Lesson. Two workloads that share code are not one workload. Let them share the code and give each the capacity model its own requirement implies.

quality-and-publication

Proving an output is good enough, and putting it in front of people.

ADR-09 · The quality gate blocks publication, and a waiver is a recorded act

Status: Accepted · Shown on views: 18, 22, 13

Can a rendition that scores below the quality floor be published, and if so by whom and with what record?

Context. At an assumed 1,800 title versions a day, nobody watches the output. Quality has to be decided by a number or it is not decided at all. But a number set by a quality engineer will occasionally refuse a title that content operations needs on air tonight, and the organisation will resolve that conflict somehow — either through a documented mechanism or through a quiet configuration change that nobody can find afterwards. The design's only real choice is whether the override exists in the open.

Decision. The gate is blocking by default: a rendition below the configured floor is not publishable. The only route past it is a recorded waiver by an authorised approver, with actor and justification in the append-only audit log. The person who sets the floor and the person who may waive it are deliberately different roles.

How it is realised on AWS. The gate decision service combines a full-reference perceptual score per rung — aggregate and worst-window, against assumed floors of 93 at the top rung, 78 at the lowest, and no two-second window below 60 — with structural checks for sync drift, loudness, black and frozen frames, missing segments and decodability across the device matrix. Failure produces a diagnosis naming rungs, windows and checks, and a proposed recipe adjustment; the retry runs under a new recipe version rather than the same one, because re-running a deterministic encode would fail identically. Waivers are rows in the audit log, and the human sampling queue is weighted towards narrow passes.

Option Verdict Reasoning
Blocking gate, recorded waivers, separated roles Chosen Makes the conflict visible and auditable rather than resolving it through an untracked configuration change.
Advisory gate with a quality report Rejected At this volume an advisory signal is an unread signal. Nothing would be refused and the floor would be decorative.
Blocking gate with no override at all Rejected Sounds rigorous and guarantees that the first genuinely urgent exception is handled by someone editing the floor, which is worse than a waiver.
Human review of every title Right elsewhere Right for a premium catalogue of a few hundred titles a year. Does not survive 1,800 versions a day.

What it buys

  • A quality failure is a blocked publish rather than a customer complaint, which is the cheapest place to find it.
  • Overrides are data: the audit log shows whether waivers cluster around one supplier, one codec or one deadline, which is how a floor gets corrected.
  • Separating who sets the floor from who may waive it keeps deadline pressure away from the bar itself.
  • A diagnosis plus a proposed recipe adjustment makes a failure actionable rather than a dead end.

What it costs

  • A premiere can be blocked at an inconvenient hour, and the escalation path has to work at that hour.
  • The floor is a judgement call that will be wrong in both directions until telemetry corrects it, and being wrong strictly is the more visible error.
  • Perceptual scoring every rung of every title is a real compute cost on top of the encode itself.
  • A waiver culture can develop, and the only defence is that waivers are counted and visible.

Choose differently when. If post-publication telemetry showed the floor refusing renditions that viewers demonstrably do not notice, the floor should move — that is the gate working, not failing. If waivers became routine rather than exceptional, the conclusion is that the floor is mis-set or the ladder solver is under-performing, not that the gate should be advisory.

Why it holds up over time. Perceptual metrics will improve and the specific floors will move with every codec generation, so the numbers in this record are the least durable part of it. The structure — a blocking numeric gate, an explicit audited override, and a separation between the person under deadline pressure and the person who sets the bar — is organisational rather than technical and should outlast several metrics.

Lesson. Every quality bar will be overridden eventually. Design the override before the first incident, name who may use it, and count it — otherwise the bar is edited instead.

ADR-10 · Playback telemetry is a second, post-publication gate

Status: Accepted · Shown on views: 09, 18, 22

Is quality assurance finished at the moment of publication, or does the platform keep judging a rendition after viewers have it?

Context. A perceptual score compares an encode to its source. It does not know that one television's decoder mishandles a particular profile, that a chipset drops frames above a certain bitrate, or that a rung is being selected far more often than the ladder anticipated. Those failures are only visible from the field, and they are the ones viewers actually experience. A pipeline whose quality story ends at publication finds out about them through customer support, by which point the title has been live for days.

Decision. Rebuffer rate, rung distribution and decode errors by device class are consumed as an inbound interface and act as a post-hoc gate that can automatically quarantine a published rendition and trigger a manifest revision. The pre-publication gate remains necessary and is explicitly not sufficient.

How it is realised on AWS. Telemetry is drawn on the integration view as inbound, beside the delivery interfaces, rather than as an export — that placement is the decision. Events land in BigQuery partitioned by publish date and joined to rendition digests, so a regression can be attributed to an encoder build or a ladder change rather than to a title. A quarantine removes the rendition from the manifests that advertise it, which is a revision rather than a withdrawal of the title; because manifests are cached for an assumed 30 seconds and segments for 30 days, the revision takes effect almost immediately and invalidates no cached byte.

Option Verdict Reasoning
Telemetry as an automatic post-publication gate Chosen Catches the device-specific and selection-pattern failures that no full-reference metric can see, and does it in hours rather than through support tickets.
Telemetry as a dashboard for the quality team Rejected Turns a gate into a report. At 1,800 versions a day a report is read about the titles someone already suspected.
Expand the pre-publication device matrix instead Deferred Worth doing and cannot be complete: the matrix tests decodability on devices we have, not selection behaviour on networks we do not.
Manual withdrawal on support escalation Rejected The current state of the art in many pipelines, and it measures quality in units of complaints.

What it buys

  • Device-specific failures are found by the platform rather than by viewers, and attributed to a build or a ladder rather than to a title.
  • A quarantine is a manifest revision, so remediation costs nothing at the edge and needs no re-encode to take effect.
  • The assumed false-pass target of 0.2% of titles per quarter becomes measurable, which makes the pre-publication floor tunable against evidence.
  • The same telemetry feeds ladder tuning, so the quality loop and the cost loop share one signal.

What it costs

  • Automatic quarantine can remove a rung from a working title on a noisy signal, so the thresholds need hysteresis and a blast radius.
  • Telemetry arrives with a delay and at a volume that costs real money to retain — an assumed 13 months at full granularity.
  • Attribution depends on the digest chain being intact end to end, so this gate is only as good as ADR-03 and ADR-04.
  • Device-class taxonomies drift as the player estate changes, and a stale taxonomy silently misattributes failures.

Choose differently when. If the player estate were narrow and stable — a single device family, say — a pre-publication matrix could plausibly be complete and the post-hoc gate would be redundant. If telemetry quality were poor enough that quarantines were mostly false positives, the gate should degrade to alerting a human rather than acting, which is a change of autonomy rather than of principle.

Why it holds up over time. The player estate gets more diverse over time, not less, and every new codec adds device-specific decode behaviour that no laboratory matrix fully predicts. The case for judging quality in the field strengthens as the estate fragments. What will change is how much autonomy the gate is given.

Lesson. A metric that compares output to input cannot see the thing that only happens on someone's television. Close the loop from the field, and give the loop the authority to act.

ADR-11 · Atomic manifest swap is the only publication primitive

Status: Accepted · Shown on views: 13, 19, 04

Does a title become available rendition by rendition as encoding finishes, or in one indivisible step?

Context. A full ladder across several codecs completes over hours. The natural implementation reveals renditions as they land, because each one is useful the moment it exists. The consequence is that for most of those hours the title is in a state nobody designed: a player may see a top rung with no fallback, or a ladder missing the rungs a phone needs, and the set of things a viewer can experience is the set of partial completions. At midnight on a premiere, that state is what the audience gets.

Decision. Publication is an atomic swap of a manifest pointer per title and territory. An incomplete ladder publishes as a narrower manifest or not at all. A launch subset may be published ahead of the full ladder, with the remainder added by later manifest revisions. Rollback is a pointer move to the previous revision, completing within an assumed 120 seconds and with no dependency on the encode fleet.

How it is realised on AWS. The publisher verifies that every segment a candidate manifest references is durably present and gate-passed before emitting it; a manifest pointing at an absent segment is treated as a publication defect rather than a cache miss. The swap updates the catalogue projection, which the origin and the catalogue service read. Because manifests are cached for an assumed 30 seconds and segments are immutable and content-addressed, a revision — adding rungs, narrowing a ladder, quarantining a rendition, withdrawing a title — propagates in seconds and invalidates nothing at the edge.

Option Verdict Reasoning
Atomic manifest swap, launch subset, revisions Chosen Every state a viewer can reach is a state somebody chose, and rollback needs no fleet.
Progressive reveal as renditions complete Rejected Cheap and makes the set of viewer-visible states equal to the set of partial completions, which nobody reviewed.
Publish only when the full ladder is complete Rejected Safe and misses the premiere: the full ladder takes an assumed four hours and the embargo is at midnight.
Per-rendition availability flags read by the player Rejected Moves the publication decision into the player, which is out of this boundary and cannot be rolled back from here.

What it buys

  • Nothing is ever half-published, and every reachable state is a deliberate one.
  • Rollback is a pointer move with no encode, no packaging and no re-upload, which is what makes the assumed 120-second target credible.
  • A launch subset lets a premiere ship on time without pretending the full ladder is ready.
  • Immutable segments plus a 30-second manifest cache mean revisions are near-instant and free at the edge.

What it costs

  • The publisher must verify durability across potentially thousands of segments before each swap, which is latency on the critical path at midnight.
  • Launch-subset width is a judgement call with a real product consequence, and on day zero there is no demand signal to inform it.
  • A narrow first manifest means some viewers genuinely get fewer rungs for a while, and that has to be acceptable.
  • Manifest revisions multiply per title and territory and device class, so the projection carries more rows than an intuitive design would.

Choose differently when. If the full ladder could be produced inside the embargo window — a much faster fleet, or a much narrower ladder — the launch subset becomes unnecessary complexity and publishing complete is simpler and better. If territories and device classes multiplied far beyond the assumption, per-title atomicity might need to become per-title-per-territory batching to keep the swap bounded, which is a change of granularity rather than of principle.

Why it holds up over time. Atomic pointer swaps over immutable content is the same pattern as a blue-green deploy, a database index swap and a content-addressed release: it has been the right answer in every domain that needed a reversible publication, and it does not depend on anything about video. Segment immutability is what makes it cheap, and immutability is not going away.

Lesson. If a long-running process can produce many intermediate states, publish an indivisible pointer instead — and the states nobody chose become unreachable rather than merely unlikely.

economics-and-lifecycle

Deciding which renditions are worth existing, and for how long.

ADR-12 · The bitrate ladder is solved per title from a complexity profile

Status: Accepted · Shown on views: 02, 10, 18

Does every title get the same ladder, a ladder solved for that title, or a ladder that varies shot by shot within the title?

Context. A fixed ladder is wrong for almost everything. Animation at 1080p needs a fraction of the bitrate of grainy 35 mm film at the same resolution and perceptual quality; sport needs more than either. A fixed ladder therefore wastes bitrate on the easy content and starves the hard content, and does both at every rung. Per-title solving corrects that at the cost of analysis compute. Per-shot solving corrects it further and introduces a harder problem: rungs must stay segment-aligned across a title so a player can switch at any segment, and per-shot variation pulls against that.

Decision. The ladder is derived per title from a complexity profile — shot boundaries, per-shot spatial and temporal complexity, grain, dominant motion, letterbox geometry, and the effective source resolution after any upscale in the master. The egress saved is weighed against the analysis compute spent, per title, and the comparison is recorded. Per-shot variation is deferred to Phase 3, with segment alignment preserved as a hard constraint.

How it is realised on AWS. The analyser runs as Cloud Run jobs before any encode is scheduled, and the solver writes its result into the recipe, so the ladder is frozen before a single chunk runs. Rungs above the effective source resolution are omitted rather than upscaled, and a rung exists only if some supported device class and some observed network condition can select it. The cost comparison is a row in BigQuery per title, which is what makes the trade arguable rather than asserted: at an assumed 40 PB/day of delivered egress, one percent of average bitrate across the ladder is worth about 400 TB/day.

Option Verdict Reasoning
Per-title ladder from a complexity profile Chosen Captures most of the available bitrate saving at a bounded analysis cost, and keeps segment alignment trivially intact.
Fixed catalogue-wide ladder Rejected Simplest, cheapest to operate, and wrong in both directions simultaneously on a catalogue spanning animation and 35 mm grain.
Per-shot ladder variation Deferred More saving and a real risk to cross-rung segment alignment. Phase 3, and only once the per-title gain has been measured.
Per-device-class ladders Rejected Multiplies renditions by device class, which multiplies the storage and cache problem ADR-01 exists to contain.

What it buys

  • Easy content stops being over-encoded and hard content stops being starved, at every rung rather than on average.
  • The ladder decision becomes answerable in money, which is what lets it be tuned rather than debated.
  • Refusing to upscale removes a whole class of rungs that cost storage and egress and deliver nothing.
  • Freezing the ladder into the recipe before any chunk runs means the whole title is encoded against one coherent decision.

What it costs

  • Analysis is compute spent before any output exists, on every title, including the ones nobody will watch.
  • Variable ladders make cross-title quality comparison harder, because two titles' rung-three are no longer the same thing.
  • The solver is a model, and a model that drifts produces ladders nobody notices are wrong until telemetry says so.
  • Per-title rung counts make cache and storage forecasting less predictable than a fixed ladder.

Choose differently when. If analysis cost per source hour approached the encode cost — a far more expensive analysis, or much cheaper encoding — the per-title gain would stop paying for itself on the long tail, and the right answer becomes per-category ladders with per-title solving reserved for high-demand titles. Conversely, if the measured saving were large and segment alignment proved tractable under per-shot variation, Phase 3 moves forward.

Why it holds up over time. Content-adaptive encoding has moved from a research result to standard practice over the last decade, and every codec generation widens the gap between a fixed ladder and a solved one because codecs get better at exploiting content structure. The specific profile features and the solver will be replaced repeatedly; deciding per title rather than per catalogue is what survives.

Lesson. Catalogue-level defaults are where the waste lives. Make the expensive decision once per item, record what it cost to make, and the trade stops being a matter of opinion.

ADR-13 · Rendition residency is an economic decision, re-made continuously

Status: Accepted · Shown on views: 19, 11, 15

On what rule is a rendition demoted to a cheaper storage class, and on what rule is it deleted outright?

Context. ADR-01 permits eviction; it does not say when. The obvious rule is age, because every object store implements it and it needs no thought. Age is also almost unrelated to the thing that matters: a ten-year-old title in a prestige collection may be watched steadily while a title from last month is watched twice. The quantity that actually decides whether a rendition should exist is the cost of keeping it against the cost of making it again, over whatever horizon the platform is willing to plan for — and both sides of that comparison move with storage prices, compute prices and the title's own demand.

Decision. A rendition is demoted and eventually evicted when its regeneration cost falls below its retention cost over the policy horizon. The rule is economic, not temporal, and is re-evaluated continuously against observed demand. Eviction is not an end state: a request for an evicted digest re-enters the lifecycle at generation.

How it is realised on AWS. A demand-tier controller on GKE reads per-digest request rates from the telemetry store and writes a residency class onto each rendition row; Cloud Storage lifecycle rules then execute the demotion and deletion that the controller has decided. The lifecycle view is drawn as a ring with seven states precisely so that Evicted has an outgoing arrow — a request re-enters at Generated through the on-demand path of ADR-08. Targets are assumed: at least 70% of catalogue hours held at the cheapest tier or not held at all, rendition storage no more than 18% of total storage spend, and on-demand generation no more than 6% of encode spend.

Option Verdict Reasoning
Economic rule on demand, regeneration cost and retention cost Chosen Compares the two quantities that actually decide the question, and re-prices itself as costs and demand move.
Age-based lifecycle rules only Rejected Free to implement and uncorrelated with demand. Evicts the steadily-watched prestige title and keeps last month's failure.
Manual curation of residency by content operations Right elsewhere Right for a few hundred titles with strong editorial opinions. Unworkable across 1.2 million rendition sets.
Keep everything resident in the cheapest class Rejected Avoids every fallback and keeps the full object count, which is a bill that scales with catalogue rather than viewing.

What it buys

  • Residency re-prices itself as storage and compute costs move, without anyone rewriting a policy.
  • The long tail costs roughly what it is worth, which is the saving ADR-01 was taken for.
  • The JIT share of spend becomes a single number that says whether the policy is tuned, visible on the observability grid.
  • A sudden spike on a cold title promotes it automatically, so popularity fixes its own residency.

What it costs

  • Aggressive eviction converts into viewer-visible fallbacks, and the feedback loop between the two is indirect and delayed.
  • The controller needs per-digest demand data at a granularity that is itself expensive to collect and retain.
  • Regeneration cost is an estimate, and a wrong estimate evicts things it should not in a way that is hard to notice.
  • Two systems now decide residency — the controller and the lifecycle rules — and a disagreement between them is a confusing class of bug.

Choose differently when. If the on-demand spend ceiling were breached persistently, the policy is too aggressive and the horizon should lengthen rather than the ceiling rise. If a rights holder's contract forbade disposing of deliverables for part of the catalogue, residency becomes a per-title contractual floor, which this design can express as a policy value — which is why it is a policy value and not an architectural constant.

Why it holds up over time. The comparison itself — keep versus remake — is timeless, and expressing it as a rule rather than a schedule means the design absorbs changes in both prices without redesign. What will change continuously is the horizon and the thresholds, which is the intended behaviour rather than a weakness.

Lesson. When the obvious policy lever is time and the real variable is money, write the rule in money. Age-based retention is a proxy nobody chose and everybody inherits.

protection-and-operations

Keeping an unreleased master and a content key safe, and recovering a region.

ADR-14 · Regional recovery is regeneration from custody, not rendition replication

Status: Accepted · Shown on views: 17, 11, 22

After the loss of the encode region, is the platform recovered by replicating renditions into a second region, or by regenerating them there from masters and recipes?

Context. This is ADR-01 meeting a disaster-recovery review, and it is where the cache model is most likely to be quietly reversed. The familiar answer is to replicate everything to a standby region and fail over, because that is what the runbook template says. Applied here it means replicating an assumed 1.2 million rendition sets, paying twice for storage and once for egress on everything, in order to protect bytes that the architecture has already established are a function of two much smaller things.

Decision. Masters, recipes, gate scores and the audit log are dual-region. Renditions are not replicated for availability — they are regenerated. The standby region carries a warm control plane and no encode fleet. After a region loss, the published manifest degrades to the surviving rung set and refills, with an assumed 30-minute p95 for a title's launch subset.

How it is realised on AWS. The deployment view shows three groupings: a primary encode region with the fleet, a dual-region custody band holding masters, recipes and audit, and a standby region with a warm control plane and a degraded origin and nothing else. Control-plane recovery is an assumed RTO of 15 minutes against an RPO of 5 seconds on Spanner. In-flight jobs restart from completed chunks, which is a direct consequence of chunk idempotence in ADR-05. Cross-region regeneration is exercised on a schedule in Phase 3, because a recovery mechanism nobody drills is a hypothesis.

Option Verdict Reasoning
Dual-region custody, regeneration on failover, no standby fleet Chosen Replicates the small irreplaceable things and buys a regeneration budget instead of a second copy of everything.
Full rendition replication to a standby region Right elsewhere Right under a pre-encode-and-preserve model, or where a contract specifies a warm second copy. Doubles storage to protect derivable bytes.
Active-active encode in both regions Rejected Doubles the fleet footprint and adds a cross-region consistency problem on job state that nothing in the requirement asks for.
Backup and restore of renditions from archive Rejected Slower than regenerating them, and pays archive storage for every rendition to do it.

What it buys

  • Disaster recovery costs the replication of masters and recipes — orders of magnitude smaller — plus a regeneration budget.
  • The standby region is cheap enough to keep genuinely warm, because it holds a control plane rather than a fleet.
  • Chunk idempotence means in-flight jobs resume rather than restart, so a failover does not lose a night's encoding.
  • The degraded state is explicit and drawn: a narrower rung set, refilling, rather than an outage.

What it costs

  • Immediately after a failover, viewers in every territory may get a narrower ladder than usual until the refill completes.
  • Recovery time now depends on encode capacity in the standby region, which is the resource that was deliberately not pre-provisioned.
  • A regeneration storm during a failover competes with the on-demand path for exactly the capacity that is scarcest.
  • The mechanism is only credible if it is drilled, which is ongoing operational cost rather than a one-off build.

Choose differently when. If a contract or a regulator required a warm second copy of deliverables rather than a demonstrated ability to reproduce them, replication becomes mandatory for that part of the catalogue. If regeneration throughput in a standby region proved unobtainable at short notice — a capacity constraint rather than a cost one — the honest fallback is replicating the launch subset only, which protects the viewer-visible case at a fraction of the full cost.

Why it holds up over time. The decision follows from ADR-01 and inherits its durability. The specific regional topology will change; the principle that you replicate the inputs to a function rather than its outputs is what persists, and it gets more valuable as the ratio between catalogue size and master size grows.

Lesson. Replicate the inputs, not the outputs. A disaster-recovery plan that copies everything derivable is usually a plan written before anyone asked what was derivable.

ADR-15 · The encode fleet is an untrusted zone, and no fallback weakens protection

Status: Accepted · Shown on views: 20, 21, 22

What is the assumed attack on this platform, and what is the one thing the design refuses to degrade under any failure?

Context. Most platforms place their trust boundary at the API and treat internal compute as trusted. That is the wrong boundary here. The encode fleet's entire job is to run a decoder over bytes supplied by third parties, and codec parsers are among the most consistently exploitable software in production anywhere. Meanwhile the assets involved are unusually sensitive in a specific way: an unreleased master is a pre-embargo copy of something whose leak is a contractual and commercial event, and a content key is the thing that makes an entire catalogue's encryption meaningful. Both of those are handled by the same fleet.

Decision. The encode fleet is treated as untrusted. Workers decode inside a sandbox with no egress, hold a short-lived workload identity granting read on one master prefix and write on one rendition prefix, and nothing else. Content keys are obtained from an external key service per packaging job, held in memory for that job, and never persisted in pipeline stores, logs, metrics or worker disk. And one invariant admits no exception: no fallback may weaken protection — quality, latency and rung availability may all degrade, encryption may not.

How it is realised on AWS. The trust-zone view gives the fleet its own zone between the perimeter and the trusted platform, which is unusual and deliberate. gVisor isolates the decoder process; egress is denied at the network level so a compromised parser has nowhere to send anything. Pre-release masters sit in the custody zone with CMEK, per-object read auditing and alertable bulk reads, with per-rights-holder key separation where a contract requires it, and pre-release renditions are served from an origin that is not publicly routable until publication. The invariant shows up concretely in the failure table: a key service outage stalls publication and that is the whole of the handling — there is no unencrypted path to fall back to.

Option Verdict Reasoning
Untrusted fleet zone, sandboxed decoders, keys never persisted, protection never degraded Chosen Places the boundary where the hostile input actually arrives, and removes the one fallback that would be reached for under pressure.
Trusted internal compute with a perimeter boundary Rejected The conventional topology, and it puts the trust boundary in the one place the attack does not come from.
Cache content keys in the packaging tier for throughput Rejected Trades a contractual exposure for a latency gain nobody asked for, and makes the pipeline a place keys live.
Allow an unencrypted path for internal review workflows Rejected Creates exactly the exception that gets used during an incident, and it would be used by someone with a deadline rather than a threat model.

What it buys

  • A codec-parser compromise is contained to a sandbox with no egress and credentials that reach two prefixes.
  • Key custody never passes to the pipeline, so a pipeline compromise does not become a catalogue-wide decryption event.
  • Per-object read auditing on masters makes a pre-embargo leak investigable rather than deniable.
  • Removing the unencrypted fallback removes the option that would otherwise be taken at 2am under commercial pressure.

What it costs

  • Sandboxing costs encode throughput, and on a fleet of this size that is a measurable share of the compute bill.
  • Per-job key fetches put the key service on the packaging critical path, so its availability becomes a publication dependency.
  • A key service outage delays a premiere, and the design offers no mitigation beyond retry — by intent.
  • Narrow per-worker credentials and no egress make debugging a failing encode genuinely harder for operators.

Choose differently when. Nothing in ordinary operation flips the protection invariant; it exists precisely to be unflippable. The realisation can change: if a provider offered hardware-level isolation for the decode step with less overhead than a sandbox, that replaces gVisor without touching the decision. If key-service availability proved to be the dominant cause of missed premieres, the right response is a more available key service, not a weaker encryption path.

Why it holds up over time. Codec parsers have been a reliable source of memory-safety vulnerabilities for as long as codecs have existed, and nothing about the current direction of either media tooling or sandboxing changes that. The sensitivity of pre-release content is contractual rather than technical and does not decay. Sandbox overheads are falling, which makes the decision cheaper over time.

Lesson. Put the trust boundary where the hostile input arrives, not where the authentication is. Then name the one thing you will never degrade, and delete the fallback that would let you.

Every package used, in one table

Ten terms that carry specific meaning in this package, several of which mean something looser in general industry use.

Package What it is What it does here Considered instead
Mezzanine The high-bitrate intermediate master delivered by a rights holder — an assumed 80 Mbps, about 36 GB per source hour — from which every rendition is derived. The only non-derivable artefact in the pipeline. Immutable on arrival, retention-locked, dual-region. Sometimes called the source, the house master or the ProRes; in this package it is always the delivered intermediate.
Recipe An immutable, content-addressed document holding the solved ladder, the analysis inputs, every encoder parameter and the pinned encoder build. The authoritative derived artefact and the only sanctioned encode input. The join key between a master and everything derived from it. Often called a job template or an encoding profile, both of which usually imply something mutable and shared.
Rendition One encoded version of a title at one rung, in one codec, produced by one recipe and one encoder build. A cache entry: evictable under an economic policy, regenerable on demand, carrying no RPO. Elsewhere a rendition is typically a durable deliverable. That reading is explicitly rejected here (ADR-01).
Rung One step of the bitrate ladder — a resolution and target bitrate pair that a player may select. The unit the ladder solver decides, the gate scores and the manifest advertises per device class. Also called a variant, a representation or a profile depending on the streaming format.
Launch subset The few rungs that carry the majority of first-day sessions, published ahead of the full ladder. What makes an assumed 45-minute premiere turnaround possible without pretending the full ladder exists. No settled industry term; sometimes described as a fast-publish or day-one ladder.
Reclaim The provider withdrawing a Spot worker mid-task — an assumed 18% hourly at steady state. The expected steady state rather than a failure. Costs at most one chunk and is drawn as a stage of the flow. Also preemption or interruption. Treating it as an error path is the mistake this design is built to avoid.
Common encryption Encrypting segments once under a scheme that multiple DRM systems can all issue licences against. What lets one encrypted segment set serve three DRM systems, so adding one is an integration rather than a re-packaging job. Sometimes conflated with DRM itself; here DRM is the external licence service and this is the encryption scheme.
Manifest The per-device-class document advertising which rungs, codecs, audio tracks and encryption schemes are available. The only mutable published artefact, cached for an assumed 30 seconds. Publication, rollback, revision and quarantine are all manifest operations. Also the playlist or the MPD. Not to be confused with the delivery manifest a supplier sends with a master.
Gate The blocking combination of a full-reference perceptual score per rung, structural checks and a device decoder matrix. Decides publishability. Passing is required; a waiver is the only route past it and is recorded with actor and reason. Often an advisory QC report. The distinction between blocking and advisory is the whole of ADR-09.
Residency Which storage class a rendition currently occupies, or whether it exists at all. Decided continuously by comparing regeneration cost against retention cost over the policy horizon, not by age. Usually expressed as a retention schedule, which is a proxy for demand that nobody actually chose.
The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.