An image-derivative pipeline runs as functions: about 270 million invocations a month at roughly 1.8 seconds and 2 GB of memory each, costing about 16000 dollars a month at 2026 per-invocation prices. The team wants to move it to a container fleet on one-year reserved capacity. Give the sequence, say where output can diverge, and name the point of no return.
Show the full answer Hide the answer
Does the move even pay
270M × 1.8 s × 2 GB is roughly 970 million GB-seconds a month. The same work as sustained compute is 270M × 1.8 s of CPU time, about 135000 CPU-hours, which over a 730-hour month is roughly 185 cores running flat out, or about 265 cores at a 70% utilisation target. At an order-of-magnitude 0.03 dollars per vCPU-hour on one-year reserved capacity in 2026, that is about 6000 dollars a month.
So the saving is roughly 10000 dollars a month, or 120000 a year, against 8 to 12 engineer-weeks to migrate and something like 0.2 of an engineer to run the fleet afterwards. It pays, but by about one engineer's salary, which is the scale at which you should also ask whether reserved capacity for the existing functions or a memory reduction gets you a third of the benefit for a week of work.
The sequence, each step reversible
- Put the work behind a queue with an explicit job contract (source key, target derivative, output location). Until the executor is swappable, nothing else is safe.
- Run containers as a second consumer at 5%, writing outputs to a shadow prefix rather than the live one.
- Compare outputs, not exit codes. Byte equality will fail: a different version of the image library changes chroma subsampling and metadata, so files differ while images look identical. Compare a perceptual hash, the pixel dimensions and the file size within a tolerance band, and review the mismatches by eye for the first week.
- Ramp 5% to 50% to 100%, keeping the function path deployed and warm the whole time.
- Run both for one full cycle after 100%, because the tail of odd inputs arrives slowly: CMYK JPEGs, animated GIFs, 20000-pixel panoramas, truncated uploads and the handful of files with broken EXIF orientation.
Where data diverges, and how you would know
Divergence is in the output bytes, and it is invisible unless you look at pixels. The second divergence is subtler: if the new encoder produces a different default quality, every derivative rendered from now on differs slightly from the millions already cached at the CDN, so the same page can show two visibly different crops of one photo. Detect it by rendering a fixed set of 500 reference inputs through both paths on every deploy and diffing the perceptual hashes.
The point of no return
Not the traffic switch, which is a weight change. It is any change to the output format, the encoder defaults or the derivative cache keys, because that invalidates cached derivatives and re-rendering the back catalogue is a batch job over the whole corpus. Freeze output format for the migration and change it, if ever, as a separate project with its own backfill.
Decommissioning the function deployment is the second one-way step, and it should lag the cutover by a month.
How long it really takes
8 to 12 weeks, and the code is the small part. The comparison harness, the odd-input tail and the on-call handover for a fleet the team has not run before dominate. Add capacity work the functions did for free: concurrency limits so 265 cores do not open 2000 connections to the metadata database, a warm pool for the morning ramp, and a queue-depth autoscaler.
When not to do this at all
Below roughly 2000 dollars a month the saving does not cover the engineer-weeks, let alone the on-call. The genuine reasons to move at any spend are the limits rather than the price: an execution-duration ceiling you keep hitting, a need for a GPU or a large local scratch disk, or a startup cost you cannot amortise. If the motivation is only the invoice, try reserved concurrency pricing and a memory-size sweep first: halving the memory on a CPU-bound function halves the bill without a migration.