# Streaming Video Encoding & Packaging Pipeline

**Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records**

Almost nobody outside the industry has a name for this system, and almost everybody has met it. A show lands at midnight and plays instantly in 4K on the television, while the same episode on a phone starts soft and sharpens eight seconds in. A film from 1974 looks worse than it should on a big screen; a cartoon from last week looks flawless at a quarter of the bitrate. A new release stalls for everyone in the first hour and is fine by morning. None of that is the network and none of it is the player — it is the encoding and packaging pipeline, the thing that turns one enormous studio master into the few hundred small files a player actually downloads, and decides per title how many of them are worth making. This package is that pipeline for a consumer video-on-demand service: an assumed 400,000 titles and 250,000 source hours, 3,500 source hours a day arriving as ~126 TB of immutable masters, 220,000 chunk-encode tasks a day across a reclaimable fleet of ~14,000 slots, and 40 PB a day of delivered egress — built on dual-region archive custody, open-source encoders on per-second-billed Spot capacity, a transactional control plane, and a warm fleet that regenerates a rendition while a viewer waits.

The hard part is not video compression. It is that the catalogue is enormous and the viewing is not: an assumed 72% of catalogue hours attract under 0.1% of total watch time. Pre-encoding a full ladder in several codecs for all of it is a storage bill paid in perpetuity for files nobody requests; not pre-encoding it means a real viewer waits. The architecture is the answer to where that line goes.

The design rests on one sentence: **the master plus the recipe is the asset, and every rendition is a cache entry.**

The decisions that carry the design:

- **The master and the recipe are the only assets.** Renditions get no RPO, no backup and no cross-region replication. Storage becomes a retention-versus-regeneration calculation per title, and disaster recovery for 1.2M rendition sets becomes replication of two much smaller things plus a regeneration budget (ADR-01).
- **The mezzanine is immutable on arrival.** Write-once, retention-locked, verified against a delivery manifest before any work is scheduled. A redelivery is a new source version requiring explicit promotion — a supplier's upload is not an editorial decision (ADR-02).
- **The recipe is the only sanctioned encode input.** Immutable, content-addressed, pinned to an encoder build. It is the one derived artefact that cannot be recomputed, so it gets custody and an address and everything else becomes a function of it (ADR-03).
- **Renditions are identified by content hash.** A digest over source version, recipe and encoder build, which makes an encoder upgrade a catalogue-wide cache miss filled lazily by demand rather than a funded re-encode of 250,000 hours (ADR-04).
- **Bound the reclaim loss; do not checkpoint.** Shot-boundary chunks sized so a withdrawn worker costs at most a chunk-minute, with chunk duration as a measured dial between loss and boundary count rather than a constant someone once chose (ADR-05).
- **Determinism is a release gate.** An encoder build that cannot replay a golden corpus byte-identically is rejected before anything else about it is measured, because idempotent retries, content addressing and lineage are all consequences of it (ADR-06).
- **Lanes with capacity floors, ordered by deadline.** A premiere cannot be starved by a backfill and a backfill cannot be starved until a rights window expires. Urgency is computed from the deadline and the work remaining, not asserted once at submission (ADR-07).
- **Batch and on-demand share code, never capacity.** One is throughput-bound on reclaimable compute, the other latency-bound with reserved headroom; mixing them gives a waiting viewer a Spot queue or gives the batch fleet an availability target it does not need (ADR-08).
- **The gate blocks, and a waiver is a recorded act.** At 1,800 title versions a day nobody watches the output, so a number decides publication — and the person under deadline pressure is deliberately not the person who sets the floor (ADR-09).
- **Playback telemetry is a second gate.** A full-reference metric cannot see what only happens on someone's television, so the field closes the loop and is given the authority to quarantine a published rendition (ADR-10).
- **Atomic manifest swap is the only publication primitive.** An incomplete ladder publishes narrower or not at all, so every state a viewer can reach is one somebody chose — and rollback is a pointer move needing no fleet (ADR-11).
- **The ladder is solved per title.** Derived from a complexity profile and weighed as egress saved against analysis compute spent, because catalogue-level defaults are where the waste lives and one percent of average bitrate is worth ~400 TB/day (ADR-12).
- **Residency is written in money, not in time.** A rendition is demoted and evicted when regeneration cost falls below retention cost over the policy horizon, and eviction is not an end state — a request re-enters the lifecycle at generation (ADR-13).
- **Replicate the inputs, not the outputs.** The standby region carries a warm control plane and no encode fleet; after a region loss the manifest degrades to the surviving rung set and refills, which is a visible consequence rather than a hidden one (ADR-14).
- **The encode fleet is untrusted, and protection never degrades.** A container escape through a codec parser is the assumed attack, not a request to the API. Quality, latency and rung availability may all degrade under failure; encryption may not, and the fallback that would let it has been deleted (ADR-15).

The architecture one-pager (including why the design should still hold up in ten years, what a prototype would have to prove, and the six risks that would change it) and the full decision record appear on the landing page of the diagram set, directly below the index of views. The same content is published as [docs/architecture-one-pager.md](docs/architecture-one-pager.md) (~15 min) and [docs/decision-record.md](docs/decision-record.md) (~63 min).

Every quantity in this package is a **stated assumption** for a large consumer video-on-demand service, invented so the architecture commits to something arguable. None is taken from a published figure for any named streaming service. The requirement it was built from is [ask.md](ask.md).

## Rebuilding

```bash
bash scripts/build.sh          # from this directory; needs Node 20+ and nothing else
```

Specs are authored as `specs/part-{a,b,c}.json`, `specs/manifest-{a,b}.json` and
`specs/adr-{onepager,records-a,records-b}.json`. The assembled `specs/views.json`,
`specs/manifest.json` and `specs/adr.json` are generated — edit the parts, never the
assembled files.
