Streaming Video Encoding & Packaging Pipeline · View 15 of 22 · Runtime
The budget is a circuit breaker
- On-demand generation is capped at an assumed 6% of total encode spend. Past the cap the viewer is served the nearest resident rung rather than held.
- If the cap is being hit routinely, the eviction policy is wrong, not the cap — that signal is on the observability grid for exactly that reason.
Two mechanisms that make it work
- Coalescing on the rendition digest: concurrent requests for the same missing rendition produce one encode.
- Archive-class master reads are the slow part, so the first-byte budget is really a restore-plus-encode budget. Assumed p95 ≤ 1.5 s, p99 ≤ 3.5 s.
Open
- Placement — batch fleet, dedicated warm pool beside the origin, or the edge — is Question 7. This view assumes the warm pool, which is the only option compatible with both the latency budget and the spend cap, but it is an assumption and not a finding.