pattern

Micro-Batching

also called Mini-Batch Streaming, Discretised Stream Processing

Executing a stream as a sequence of small bounded jobs on a fixed trigger interval, which buys batch failure semantics and atomic sink commits at the price of a latency floor equal to the interval.

sparkmicro-batchlatencytriggerexactly-once

Two teams describe their pipeline as streaming. One processes each record as it arrives with per-record state updates. The other wakes every 30 seconds, reads whatever accumulated, runs a small job over it and commits. Both are called streaming, and only one of them can tell you that a given 30-second window either landed completely or did not land at all.

Micro-batching is the second shape. Zaharia and colleagues formalised it in 2013 as discretised streams: an unbounded input is cut into bounded batches, and each batch is an ordinary fault-tolerant job. Spark Structured Streaming's default execution mode works this way, with a trigger interval choosing the cut.

The appeal is not throughput. It is that every recovery question has a batch answer. A batch either committed or it did not, so retry is idempotent at the batch level, and the sink can be written inside one transaction per batch rather than per record.

Why it matters

Teams reach for continuous processing because "real time" sounds like the stronger choice, then discover that per-record exactly-once requires a transactional sink, a checkpoint store, a watermark policy and a late-data decision. Micro-batching gives a weaker latency guarantee and a much simpler operational story, and for the large majority of analytical pipelines the latency difference is below the threshold anybody can act on.

The second reason it matters is cost shape. Per-trigger overhead is roughly constant, so the economics of the interval are not linear.

Implementation patterns

  • Pick the trigger from the decision latency, not from how fresh the data could be. A dashboard read twice an hour does not need a 5-second trigger.
  • One commit per batch to a transactional sink, with the batch id recorded in the destination so a retried batch is recognised and skipped.
  • Checkpoint the offsets per batch, so recovery means "re-run batch 8,412" rather than "restore 200 GB of state".
  • Watch batch duration against trigger interval as the primary health signal. Once duration exceeds the interval, batches queue and the lag grows without bound.
  • Use a trigger of "available now" for backfills, which runs the same code over everything retained and then stops, so the backfill and the live job are one implementation.

Industry example

Spark's discretised-stream design (2013 paper, then Structured Streaming from 2016) is the documented reference, and its trigger modes make the trade explicit in configuration rather than in architecture: a fixed interval, a one-shot run, or continuous processing with per-record latency and a narrower feature set.

The pattern recurs outside Spark. A quick-commerce marketplace running inventory projections every 60 seconds, a telemetry pipeline flushing to a columnar store every 5 minutes, and a feature pipeline refreshing daily aggregates every 10 minutes are all micro-batch systems in production, usually described as streaming because that is the word available.

Failure scenarios

  • Batch duration overtakes the interval. Batches queue, the input lag climbs linearly, and the job never recovers on its own because each batch is now larger than the last. The only recoveries are more capacity or a longer interval.
  • Interval too short for the overhead. At a 1-second trigger, planning and listing can dominate useful work, and the cluster runs at high utilisation producing very little. Cost per record rises by an order of magnitude between a 30-second and a 1-second trigger on the same volume.
  • Small-file explosion. Each batch writes at least one file per partition per sink, so a 10-second trigger over 200 partitions produces about 1.7 million files a day, and the query side slows until compaction catches up.
  • A late event that misses its batch is not handled by the batch boundary at all; it still needs a watermark and a late-data policy, which teams assume micro-batching has solved for them.

Trade-offs

Choose Gains Pays
Micro-batch at 30 s to 5 min Batch retry semantics, one sink transaction per batch, simple recovery Latency floor at the interval, p99 near twice it
Micro-batch under 5 s Lower latency with the same semantics Per-trigger overhead starts to dominate, small-file pressure
Continuous processing Sub-second latency Per-record state, checkpoint tuning, transactional sink per record

When not to use it

When something downstream acts automatically in under a second, micro-batching is the wrong shape and no amount of interval tuning fixes it: fraud declines, bidding, device control and trading all fall here.

At the other end, when the interval you actually need is an hour or more, stop micro-batching and run a scheduled batch job. A scheduled job gives you overwrite-by-partition idempotency, no always-on cluster and a trivial backfill, and an always-on micro-batch job triggering hourly is a cron job with a cluster attached to it.

Interview question

Q: A team asks to reduce their Spark Structured Streaming trigger from 60 seconds to 1 second to make a dashboard feel live. What do you ask, and what do you expect the cost to do?

What a strong answer covers: ask who acts on the dashboard and within what time, because the answer is usually "a human looks at it a few times an hour" and 60 seconds already exceeds the requirement. Then the mechanics: per-trigger overhead runs 60 times more often, output files multiply by 60 with a matching compaction and query-planning cost downstream, and the cluster must now be sized for peak per-batch work rather than average. Expect compute cost to rise several-fold and the dashboard to feel the same, because the browser's own refresh and the human's attention are the real latency. A strong answer offers the cheaper alternative: keep the trigger and make the dashboard show the data's timestamp, which addresses the actual complaint.

Quick check

Quiz: What is the best-case end-to-end latency of a micro-batch job with a 30-second trigger? About 30 seconds, since an event arriving just after a trigger waits a full interval before it is read, and the batch duration is added on top.

Flashcard: Why does halving the trigger interval more than halve the cost per record? Because each trigger pays roughly fixed planning, scheduling and commit overhead, so shrinking the interval multiplies the overhead without reducing the useful work.