concept

Watermark

The stream processor's assertion that no further events older than a given event time are expected, which is what allows a window to close.

streamingevent-timewindowing

Every stream processor faces the same unanswerable question: an hourly window covering 09:00 to 10:00 — when can the result be emitted? Waiting forever is correct and useless; emitting immediately is fast and wrong, because a mobile device that was offline will deliver its 09:45 event at 11:20.

A watermark is the system's estimate, usually derived by tracking the maximum event time seen and subtracting an allowed lateness. When the watermark passes 10:00, the window fires.

The trade-off it encodes is direct and unavoidable: a longer lateness allowance means more correct results and higher latency; a shorter one is faster and drops more late data. There is no setting that avoids the choice, and the right value comes from measuring the actual distribution of event lateness in your data rather than from a default.

What to do with events arriving after the watermark is a separate design decision, and all three answers are legitimate in different contexts: drop them and account for the loss, emit a correction that downstream consumers must handle, or route them to a side output for reconciliation. Systems that silently drop late data and never measure how much is being lost are the common case, and the loss surfaces later as a mismatch against the batch system.