Architecture One-Pager
Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records
Streaming Video Encoding & Packaging Pipeline · Solution Architecture v1.0 · Google Cloud · Media Platform Architecture · 2026-10 · 22 views · 15 architecture decision records
The master plus the recipe is the asset. Every rendition is a cache entry. Everything else in this architecture is a consequence of that sentence.
Almost nobody outside the industry has a name for this system, and almost everybody has met it. A show lands at midnight and plays instantly in 4K on the television, while the same episode on a phone starts soft and sharpens eight seconds in. A film from 1974 looks worse than it should on a big screen; a cartoon from last week looks flawless at a quarter of the bitrate. A new release stalls for everyone in the first hour and is fine by morning. None of that is the network and none of it is the player. It is the encoding and packaging pipeline, and the reason it is hard is not video compression. It is that the catalogue is enormous and the viewing is not: an assumed 400,000 titles and 250,000 source hours, of which an assumed 72% of catalogue hours attract under 0.1% of total watch time. Pre-encoding a full ladder of a dozen renditions in several codecs for every one of them is a storage bill paid, in perpetuity, for files nobody requests. Not pre-encoding them means a real viewer waits. The architecture is the answer to where that line goes.
Seven stages, two execution fleets and three storage zones. A rights holder delivers a mezzanine master; the platform verifies it against a delivery manifest, probes it, classifies it against a conformance policy, and lands it write-once in dual-region archive storage. It then analyses the content — shots, spatial and temporal complexity, grain, motion, effective resolution after any upscale in the master — and solves a bitrate ladder for that title, recording the ladder, every encoder parameter and the encoder build identifier as a single immutable, content-addressed recipe. The recipe, not the rendition, is the authoritative derived artefact. A control plane on Spanner materialises the recipe into chunk tasks on shot boundaries and dispatches them to a Spot fleet whose workers are expected to be withdrawn continuously; a reclaim costs one chunk and is rescheduled without an operator. Stitched renditions pass a blocking gate — a full-reference perceptual score per rung, structural checks, and decodability on a device matrix — before a packager segments and encrypts them once under common encryption, with content keys held in memory for the duration of one packaging job and never persisted. Publication is an atomic swap of a manifest pointer, launch subset first. After publication, a lifecycle controller demotes and evicts renditions on an economic rule, and a separate warm fleet regenerates an evicted rendition when a viewer asks for it, falling back a rung rather than making anyone wait past the budget.
What it is, and what it is not
- A pipeline that makes a master playable on every supported device — not A video compression research project
- A system whose authoritative artefacts are a master and a recipe — not A rendition library that must be preserved
- An economic argument about which renditions are worth existing — not A quality-at-any-cost encoder farm
- A platform that proves quality with a number before publishing — not A workflow that relies on someone watching the output
- Designed around compute being taken away continuously — not A highly available encode fleet
- A publisher of immutable segments and mutable manifests — not A CDN or an edge cache
- A consumer of a DRM licence service — not A DRM or licence-issuing system
- Scoped to video-on-demand — not A live or contribution encoding path
- A multiplexer of audio and subtitle tracks delivered to it — not A subtitle authoring or translation tool
The decisions that are the architecture
- The master and the recipe are the only assets (ADR-01) — Renditions get no backup, no cross-region replication and no RPO. Storage cost becomes a retention-versus-regeneration calculation per title, and disaster recovery becomes replication of the two small things plus a regeneration budget.
- The mezzanine is immutable on arrival (ADR-02) — Write-once, retention-locked, never modified in place. A redelivery is a new source version requiring explicit promotion, not an overwrite — so a supplier cannot silently change what was published.
- The recipe is the only sanctioned encode input (ADR-03) — Immutable, content-addressed, pinned to an encoder build. It is the join key between a master and everything derived from it, and the reason an encode can be re-run identically in eighteen months.
- Renditions are identified by content hash (ADR-04) — A digest over source version, recipe and encoder build. An encoder upgrade therefore becomes a catalogue-wide cache miss filled lazily by demand, rather than a scheduled re-encode of 250,000 hours.
- Bound reclaim loss rather than checkpoint (ADR-05) — Shot-boundary chunks sized so a withdrawn worker costs at most one chunk-minute, with chunk duration as tunable policy. No encoder state is serialised or revived.
- Determinism is a release gate (ADR-06) — An encoder build that cannot replay a golden corpus byte-identically is rejected. Determinism is what makes retries free, caches trustworthy and lineage true rather than aspirational.
- Lanes with capacity floors, ordered by deadline (ADR-07) — A premiere cannot be starved by a backfill, and a backfill cannot be starved indefinitely. Priority is derived from the deadline and the work remaining, not from a static field that is wrong by the second day.
- Batch and on-demand share code, never capacity (ADR-08) — One is throughput-bound on reclaimable compute; the other is latency-bound with reserved headroom. Mixing them would either give a waiting viewer a Spot queue or give the batch fleet an availability requirement it does not need.
- The quality gate blocks, and waivers are on the record (ADR-09) — A rendition below the floor is not publishable. The only route past it is a recorded waiver by an authorised approver, and the person under deadline pressure is not the person who sets the floor.
- Playback telemetry is a second gate (ADR-10) — A rendition that scores well and plays badly on a real device is quarantined after publication and the manifest revised. Pre-publication metrics are necessary and not sufficient.
- Atomic manifest swap is the only publication primitive (ADR-11) — An incomplete ladder publishes as a narrower manifest or not at all. Rollback is a pointer move completing in an assumed 120 seconds with no dependency on the encode fleet.
- The ladder is solved per title (ADR-12) — Derived from a complexity profile, with the egress saved weighed against the analysis compute spent and the comparison recorded. At an assumed 40 PB/day of delivered egress, one percent of average bitrate is worth about 400 TB/day.
- Residency is an economic decision, re-made continuously (ADR-13) — A rendition is demoted and eventually evicted when its regeneration cost falls below its retention cost over the policy horizon. Eviction is not an end state: a request re-enters the lifecycle at generation.
- Recovery is regeneration, not replication (ADR-14) — The standby region carries a warm control plane and no encode fleet. After a region loss the published manifest degrades to the surviving rung set and refills — a visible consequence rather than a hidden one.
- The encode fleet is an untrusted zone (ADR-15) — A container escape through a codec parser handling a hostile master is the assumed attack. Workers decode sandboxed with no egress, read one prefix and write one, and content keys never persist in the pipeline.
Why this should still hold up in ten years
Codecs, encoders and cloud price lists will all change inside the life of this design. The decisions above were chosen to be the ones that do not.
- The long tail is a structural property of catalogues, not a fact about 2026. Viewing concentrates on a small fraction of any large catalogue, and catalogues grow faster than viewing hours per subscriber. The argument for treating renditions as a cache gets stronger as the catalogue grows, not weaker.
- A new codec is a recipe policy change, not a migration. Because identity is a hash over source, recipe and encoder build, adopting a codec means new recipes and lazy regeneration. The design that pre-encodes everything has to schedule and fund a catalogue sweep for each new codec; this one does not.
- Reclaimable compute is getting cheaper relative to reserved compute, not dearer. Every provider prices interruptible capacity below committed capacity because it monetises idle fleet. A design whose resilience comes from task granularity rather than worker protection keeps collecting that discount as the gap widens.
- Determinism and content addressing are what let the system be re-reasoned about later. The expensive failure mode in a ten-year-old media pipeline is not a bad encode; it is a published byte nobody can explain. Recording the recipe and the build digest is cheap now and is the only thing that makes an audit, a regression hunt or a regeneration possible then.
- The one invariant that must not be traded is protection. Quality, latency and rung availability are all allowed to degrade under failure, and every failure class in this design degrades one of them. Nothing degrades encryption, because that is the trade whose cost arrives as a contract breach rather than a support ticket.
Non-functional targets
Every target below is a stated assumption, invented for a large consumer video-on-demand service so that the architecture has something to be wrong about. The view column points at the diagram where the mechanism that delivers it is drawn.
| Quality | Target | How it is met | View |
|---|---|---|---|
| Control-plane availability | ≥ 99.95% / month | Regional Cloud Run and GKE across three zones; Spanner regional; no dependency on the encode fleet | 17 |
| Publish path availability | ≥ 99.9% / month | Manifest swap is a pointer move in the catalogue projection, independent of packaging and encode | 13 |
| Rollback time | ≤ 120 s | Previous manifest revision retained; rollback is a pointer move with no fleet dependency | 19 |
| Encode fleet availability | none, by design | Spot MIGs; up to 60% withdrawal in 10 minutes absorbed by chunk granularity | 14 |
| Catalogue-lane turnaround | p95 ≤ 0.80× duration | Chunked parallel encode across ~14,000 slots with per-lane capacity floors | 14 |
| Premiere launch subset | ≤ 45 min for 120 min | Three rungs, baseline codec, premiere lane floor; ≥ 160× real-time aggregate | 04 |
| Premiere full ladder | p95 ≤ 4 h from ingest | Remaining rungs added by later manifest revisions | 19 |
| On-demand first byte | p95 ≤ 1.5 s, p99 ≤ 3.5 s | Warm GKE pool, request coalescing, archive range read; nearest-rung fallback past budget | 15 |
| Master ingest volume | ~126 TB/day | 3,500 source hours/day at an assumed 80 Mbps mezzanine; write-once dual-region archive | 10 |
| Chunk task throughput | 220 k/day, 600 k peak | Pub/Sub dispatch with Spanner task claim; no shared mutable state on the encode path | 14 |
| Delivered egress | ~40 PB/day | Per-title ladder; 1% average bitrate ≈ 400 TB/day of egress | 12 |
| Top-rung quality | ≥ 93 / 100 aggregate | Full-reference perceptual score per rung against the source, blocking | 18 |
| Worst-window quality | no 2 s window < 60 | Worst-window score stored with the rendition, not just the aggregate | 12 |
| Gate false-pass rate | ≤ 0.2% / quarter | Post-publication telemetry as a second gate with automatic quarantine | 22 |
| Master and recipe durability | RPO 0, 7-year lock | Dual-region archive buckets with object retention locks; recipes dual-region | 11 |
| Control state recovery | RPO ≤ 5 s, RTO ≤ 15 min | Spanner with a warm standby control plane in the second region | 17 |
| Rendition recovery | no RPO; refill 30 min p95 | Deliberate: loss is a regeneration from master plus recipe | 11 |
| Full-ladder encode cost | ≤ $1.10 / source hour | Open-source encoders on per-second-billed Spot capacity | 14 |
| On-demand share of spend | ≤ 6% of encode spend | Budget check at admission; exceeding it re-tunes eviction rather than raising the cap | 15 |
Scope
In scope
- Delivery intake by resumable upload, supplier pull or object-store handoff, with integrity verification against a delivery manifest before any work is scheduled.
- Deep technical probing and conformance classification — accepted, accepted-with-remediation, or rejected — with the applied remediations recorded against the source version.
- Content complexity analysis and per-title bitrate ladder solving, including codec coverage as a per-title policy and refusal to upscale above the effective source resolution.
- The recipe plane: immutable, content-addressed, encoder-build-pinned recipes as the only sanctioned encode input, with human overrides recorded with actor and reason.
- Chunked parallel encode on reclaimable capacity with deterministic, idempotent chunk tasks, lane scheduling with capacity floors, quotas and per-title spend ceilings.
- Stitching with boundary verification, and segment alignment across every rung of a title so a player may switch rungs at any segment.
- A blocking quality gate: full-reference perceptual scoring per rung, structural checks, decodability on a device matrix, human sampling, and recorded waivers.
- Packaging to segmented adaptive formats from one set of elementary streams, common encryption with keys from an external key service, and per-device-class manifests.
- Atomic manifest publication, launch subsets, rollback, takedown with key revocation, and the demand-tiered residency, eviction and regeneration of every rendition.
- Lineage from master version through recipe digest and encoder build to published byte, plus the append-only audit trail of overrides, waivers, publishes and rollbacks.
Explicitly out of scope
- The CDN and edge cache — an adjacent use case in this practice.
- The player, its adaptive bitrate logic and its device-side decoding.
- Subtitle and audio-description authoring, translation and timing.
- Catalogue, artwork and recommendation metadata systems.
- Live capture, contribution encoding and near-live fast-turnaround paths.
- The rights and licensing system that decides what may be published where.
What a prototype would have to prove
The decisions in this record are falsifiable, and most of them are falsifiable cheaply. A prototype on a few hundred source hours and a few hundred Spot cores settles the questions that would otherwise be discovered after thirty diagrams have been drawn and a year has been spent. These are the six things it should measure, and the seven scenarios it should survive.
- Determinism in practice: that the chosen encoder, wrapped and pinned, produces byte-identical output for the same source, recipe and build across machine types, kernel versions and restarts. If it does not, ADR-04, ADR-05 and ADR-06 all change.
- Boundary invisibility: the chunk duration at which stitched output stops showing quality pumping at joins, measured perceptually rather than asserted — this sets the floor for the reclaim-loss bound in ADR-05.
- Ladder value: the delivered bitrate saved by per-title solving against the analysis compute spent, on a corpus spanning grainy film, animation and sport, so the trade in ADR-12 has a measured exchange rate.
- Regeneration latency: the real archive-restore-plus-encode time for one segment of a top rung, which is the entire basis of the on-demand budget in ADR-08 and ADR-13.
- Reclaim behaviour: the actual withdrawal distribution for the chosen Spot machine family in the chosen region, because the chunk-size calculation is only as good as the interval it assumes.
- Gate agreement: how often the perceptual floor and a panel of human reviewers disagree, in both directions, which is what tells you whether the floor in ADR-09 is set anywhere near the right place.
- Sixty percent of the fleet is withdrawn in ten minutes during a premiere encode. The job's progress must stay monotonic, no chunk may be lost twice, and the lane's p95 must hold.
- A supplier redelivers a master for an already-published title. The published manifest must not change until someone explicitly promotes the new source version.
- A new encoder build is promoted. No existing rendition may be re-encoded, no existing recipe may change, and the first new recipe must pin the new build.
- A master is crafted to hang the decoder. The chunk must time out, retries must be bounded, the title must quarantine with a diagnosis, and the lane must keep draining.
- The key service becomes unavailable mid-packaging. Publication must stall and retry, already-published content must be unaffected, and nothing unencrypted may reach any public path.
- A long-tail title with every rung evicted becomes popular in one hour. Concurrent requests for the same missing rendition must produce one encode, and nobody may wait past the budget.
- The encode region is lost. The control plane must fail over inside the RTO, in-flight jobs must restart from completed chunks, and the published manifest must degrade to the surviving rung set rather than break.
Open risks, carried rather than hidden
| Risk | If it lands | Response |
|---|---|---|
| The demand prediction for a launch subset is wrong on day zero | Premiere viewers get a lower rung in the first minutes — exactly the audience least willing to forgive it, and the one most likely to say so publicly. | Keep the launch subset deliberately wider than the prediction suggests for titles above a marketing threshold, and treat the subset's width as a tunable with a cost attached rather than a model output. Question 3 stays open. |
| The chosen encoder turns out not to be deterministic across machine types | Idempotent retries, content addressing and lineage all weaken at once; a duplicate execution stops being free and the cache stops being trustworthy. | Make determinism a build-promotion gate before anything depends on it (ADR-06), and pin machine family as well as image digest if the gate cannot otherwise be met. |
| On-demand generation exceeds its spend ceiling routinely | Either the cap is hit and viewers see fallbacks often, or the cap is raised and the economic case for eviction quietly disappears. | Treat a sustained breach as a signal that the eviction policy is mis-tuned, not that the cap is wrong; the JIT share of spend is a first-class operational metric for that reason (view 18). |
| The perceptual gate passes renditions that real devices play badly | Quality failures reach viewers, and the first people to notice are the audience rather than the platform. | Post-publication telemetry as a second gate with automatic quarantine (ADR-10), and human sampling weighted towards narrow passes and new encoder builds. |
| A rights holder's contract forbids treating renditions as disposable | The critical design decision does not apply to part of the catalogue, and a second, pre-encoded regime appears alongside the first. | Make residency a per-title policy from the start so a contractual floor is a policy value rather than an architectural exception, and price the exception explicitly. |
| Archive-class restore latency makes the on-demand budget unreachable | The cache model's cost advantage survives but its viewer-facing promise does not, and the fallback becomes the normal path rather than the exception. | Measure restore latency in the prototype before committing; keep a hot copy of the master for titles above a demand threshold, which is a cheaper exception than keeping every rendition. |
The reasoning behind every component and technology choice is in the Architecture Decision Record: 15 records across 5 areas, each with the alternatives that lost and what the choice costs.