beginner 2 min answer

A streaming job handles about 400 events per second in business hours and under 5 per second overnight. The team expected the bill to track the volume. It is flat at roughly 55 dollars a day whatever the traffic does. Why is it flat and what is the cheaper shape if a 10-minute delay is acceptable?

costmicro-batchingalways-onduty-cyclecheckpoints
Show the full answer Hide the answer

What is being tested

Whether you can separate the cost of work from the cost of readiness. A streaming job is mostly the second one, and nobody discovers this from the architecture diagram - they discover it from the invoice.

The mechanism

Four of the job's cost drivers do not move with event rate.

  • Reserved parallelism. Task slots are provisioned for the peak, held for the trough. Six workers at 4 vCPU each is 24 vCPU-days a day whether they process 11 million events or 400 thousand.
  • Resident state. Keyed state has to be in memory or on local disk to answer the next event in milliseconds, so the machines cannot be released overnight without discarding it.
  • Per-interval checkpointing. A 30-second interval is 2880 checkpoints a day regardless of traffic. Overnight those are writes of nearly unchanged state, billed as object-storage requests plus the bytes.
  • Log durability. The topic holds its retention at replication factor 3 whether anyone reads it fast or slowly.

The useful summary is a duty cycle. A job genuinely busy 8 hours a day pays roughly three times what the work costs, and the ratio gets worse the spikier the traffic.

The cheaper shape

A scheduled or triggered micro-batch. Every 10 minutes: start, read from the last committed offsets to the current end, write the result, commit, exit. You pay compute for the minutes you run, which for this job is a small fraction of the day, and the failure model is a batch job's - rerun it.

What that costs: a latency floor of the interval plus startup (a cold executor pool is tens of seconds, so the real freshness is 10-11 minutes, not 10), and no place to keep large keyed state between runs - it has to live in a store the job reloads, and the reload is charged on every run.

Decision rule: if the freshness requirement is coarser than a few minutes and the state either fits a cheap reload or is not needed, the scheduled shape wins by about the duty-cycle ratio. It flips when the state is large enough that reloading dominates the run, when a decision depends on seconds-old data, or when startup is a large share of the interval.

When not to move it

Do not convert a job holding tens of gigabytes of keyed state. The reload per run erases the saving and introduces a new failure - a run that times out before it has finished loading, which then never produces output and never alarms on lag.

Also check what else depends on the job running. If the always-on job is what keeps a serving cache warm, the scheduled version hands you a cold read path every 10 minutes, and the latency regression will cost more than the compute saved.