advanced 2 min answer

At 02:14 on a Saturday a storage-triggered function starts writing its output back into the prefix that triggers it. By Monday 09:00 the account has logged 290 million invocations against a monthly forecast of about $700. The budget alert fired on Sunday at 18:40. Which design decisions made this possible?

cost-governanceserverlessguardrailsquotasthresholds
Show the full answer Hide the answer

The trigger

A code change set the output prefix to the same prefix the event notification watches. Each invocation wrote an object, each object raised an event, each event started an invocation. The recursion was invisible in review because the loop does not exist in the code; it exists in the wiring between two resources that two different modules own.

Why it propagated

Nothing bounded concurrency. The function's reserved concurrency was unset, so it scaled to the account limit and stayed there for 55 hours. The expensive line is not the one people expect: at 290 million invocations the request and duration charges land near $550, while 290 million object writes at about $0.005 per 1000 are roughly $1450, so the storage requests cost nearly three times the compute. A runaway in a serverless pipeline is usually billed by the resources it touches rather than by the function itself.

Why detection lagged

The budget alert was configured against accrued monthly spend, which arrives in billing data hours after the fact, and it fired when the threshold was crossed rather than when the rate changed. A detective control cannot react faster than the telemetry it reads, so by the time it fired the loop had been running for 40 hours. Nobody watched invocation rate, which was available in seconds.

The structural fix versus the tempting local fix

The tempting fix is to correct the prefix. It removes this instance and leaves the class.

  1. Reserved concurrency on every event-driven function, set to a few times expected peak. This is the only control in the list with no lag, because it is enforced at admission rather than observed after billing.
  2. Separate buckets for input and output, enforced in the module so the wiring cannot express the loop.
  3. Alert on rate from service metrics, not on cumulative spend from billing: invocations per minute above a multiple of the trailing baseline, and new-resource creation rate.
  4. Know what the provider catches. Lambda has detected and stopped recursive loops since July 2023, dropping requests after 16 recursive invocations, using an X-Ray lineage header. At launch it covered SQS, SNS and direct invokes. A loop through object storage was outside that coverage, which is precisely why the platform control cannot replace the concurrency limit.

Common weak answers

  • "Lower the budget threshold." A threshold on a lagging signal is still lagging. It changes when you learn, by hours, not whether the money is spent.
  • "Add an approval step for new functions." It would not have caught this: the function was already approved, and a configuration value changed.
  • "Turn on anomaly detection." Useful, and it runs on the same daily billing data. Name the signal, the threshold and the automatic action, or the control is a report.