advanced 3 min answer

Google's ASPLOS 2021 paper on its video coding unit reports 20x to 33x better efficiency than its prior well-tuned software transcoding system, for workloads including YouTube. Your platform spends about $14M a year on transcode compute and someone proposes a custom accelerator programme. Estimate the annual spend at which it pays back and say which assumption dominates the error.

youtubecustom-siliconacceleratorsestimationnreamdahl
Show the full answer Hide the answer

The assumptions, stated

Four numbers decide this, and only one of them is published.

  • Efficiency gain g: 20x to 33x, from the 2021 paper. Treat the midpoint 25x as the planning value.
  • Coverage fraction f: the share of transcode cycles the accelerator can actually run. Fixed-function silicon supports the codecs, resolutions and modes it was taped out for; everything else stays in software. Assume 70% as a base case.
  • Non-recurring engineering NRE: design, verification, tape-out, board, firmware and the team, before one frame is served. Assume $60M over three years, amortised over a five-year life. This is my assumption, not a published figure.
  • Sustaining cost: firmware, toolchain, re-spins, the team you must keep. Assume $8M a year.

The arithmetic, step by step

Annual gross saving on the covered work:

C × f × (1 − 1/g) = C × 0.70 × 0.96 = 0.67 × C

Annual cost of the programme:

NRE / 5 + sustain = 60/5 + 8 = $20M a year

Break-even spend: 0.67 × C = 20, so C ≈ $30M a year. At $14M you are short by a factor of two, and the answer is no.

The range, and which assumption dominates

Push the inputs to their plausible edges:

f g Break-even C
0.9 33x ~$23M
0.7 25x ~$30M
0.5 20x ~$42M

The 20x-versus-33x argument moves the threshold by about 2%, because once g is above 20 the term 1 − 1/g is already 0.95 to 0.97. Coverage fraction moves it by a factor of 1.8, and halving the NRE assumption moves it from $30M to about $21M. So the two numbers worth a week of research are how much of your workload the silicon can run and what the programme really costs — not the headline multiplier everyone quotes.

This is Amdahl's law wearing a purchase order. An accelerator that is 25x faster on 50% of the work caps the total gain at about 1.9x, whatever the datasheet says.

What the number rules in and out

Ruled out: custom silicon below roughly $25M a year of single-kernel spend, which is the threshold even on the optimistic edge of the table. Ruled in first, because their NRE is zero: off-the-shelf transcode accelerator cards and GPU encode blocks, a codec mix change, and cutting the output ladder. Exhaust every zero-NRE option before anyone draws a floorplan.

When this is the wrong answer

Three conditions flip it even at lower spend. The workload must stay stable for the five years the silicon lives — a codec transition mid-life strands the investment. The organisation must be able to keep a hardware team; a programme that loses its team at year two owns an un-maintainable fleet. And if the constraint is power or rack density rather than dollars, perf-per-watt can justify silicon that fails the pure cost test.

Common weak answers

  • "20x efficiency means 20x cheaper." It means 20x on the covered portion, less the programme cost. The net at $14M of spend is a loss.
  • "Build it because the big platforms did." They did it at a far larger single-kernel spend, with a chip team already on staff and a workload stable enough to tape out against.
  • Quoting a single break-even number. Give the range and name the dominant assumption, or the model is a guess with a decimal point.