Google's ASPLOS 2021 paper on its video coding unit reports 20x to 33x better efficiency than its prior well-tuned software transcoding system, for workloads including YouTube. Your platform spends about $14M a year on transcode compute and someone proposes a custom accelerator programme. Estimate the annual spend at which it pays back and say which assumption dominates the error.
Show the full answer Hide the answer
The assumptions, stated
Four numbers decide this, and only one of them is published.
- Efficiency gain
g: 20x to 33x, from the 2021 paper. Treat the midpoint 25x as the planning value. - Coverage fraction
f: the share of transcode cycles the accelerator can actually run. Fixed-function silicon supports the codecs, resolutions and modes it was taped out for; everything else stays in software. Assume 70% as a base case. - Non-recurring engineering
NRE: design, verification, tape-out, board, firmware and the team, before one frame is served. Assume $60M over three years, amortised over a five-year life. This is my assumption, not a published figure. - Sustaining cost: firmware, toolchain, re-spins, the team you must keep. Assume $8M a year.
The arithmetic, step by step
Annual gross saving on the covered work:
C × f × (1 − 1/g) = C × 0.70 × 0.96 = 0.67 × C
Annual cost of the programme:
NRE / 5 + sustain = 60/5 + 8 = $20M a year
Break-even spend: 0.67 × C = 20, so C ≈ $30M a year. At $14M you are short by a factor of two, and the answer is no.
The range, and which assumption dominates
Push the inputs to their plausible edges:
| f | g | Break-even C |
|---|---|---|
| 0.9 | 33x | ~$23M |
| 0.7 | 25x | ~$30M |
| 0.5 | 20x | ~$42M |
The 20x-versus-33x argument moves the threshold by about 2%, because once g is above 20 the term 1 − 1/g is already 0.95 to 0.97. Coverage fraction moves it by a factor of 1.8, and halving the NRE assumption moves it from $30M to about $21M. So the two numbers worth a week of research are how much of your workload the silicon can run and what the programme really costs — not the headline multiplier everyone quotes.
This is Amdahl's law wearing a purchase order. An accelerator that is 25x faster on 50% of the work caps the total gain at about 1.9x, whatever the datasheet says.
What the number rules in and out
Ruled out: custom silicon below roughly $25M a year of single-kernel spend, which is the threshold even on the optimistic edge of the table. Ruled in first, because their NRE is zero: off-the-shelf transcode accelerator cards and GPU encode blocks, a codec mix change, and cutting the output ladder. Exhaust every zero-NRE option before anyone draws a floorplan.
When this is the wrong answer
Three conditions flip it even at lower spend. The workload must stay stable for the five years the silicon lives — a codec transition mid-life strands the investment. The organisation must be able to keep a hardware team; a programme that loses its team at year two owns an un-maintainable fleet. And if the constraint is power or rack density rather than dollars, perf-per-watt can justify silicon that fails the pure cost test.
Common weak answers
- "20x efficiency means 20x cheaper." It means 20x on the covered portion, less the programme cost. The net at $14M of spend is a loss.
- "Build it because the big platforms did." They did it at a far larger single-kernel spend, with a chip team already on staff and a workload stable enough to tape out against.
- Quoting a single break-even number. Give the range and name the dominant assumption, or the model is a guess with a decimal point.