A marketplace publishes a merchant sales data product with a 15-minute freshness SLO to roughly 200000 merchant consumers. A flash sale of the kind Flipkart runs takes order volume to about 20 times normal for six hours. No pipeline fails and no alert fires. What happens to the data product over those six hours and what stops it?
Show the full answer Hide the answer
Minute by minute, what happens
Ingestion keeps up, because appending events scales with partitions and nobody added merchants. The aggregation step is where it breaks: its cost is per merchant per window, and at 20 times volume a window job that took 4 minutes takes 40 or more. From roughly minute 16 the product is out of SLO, and because each run still succeeds, nothing says so.
Runs then start overlapping. Either the scheduler skips the next window, and the gap widens monotonically, or it runs concurrently and the two writers produce small files that the query layer must merge, so read latency climbs at the same time.
Read load rises with it. Merchants watching a flash sale refresh, so consumer QPS moves several times above baseline during exactly the interval when the compute is already saturated.
Where it amplifies
The consumer side has no notion of staleness, so merchants interpret a flat number as flat sales and reprice. Support tickets arrive as "sales have stopped", which routes the incident to the commerce team rather than to the data platform, costing an hour before anyone looks at watermarks. By hour three the product is roughly 90 minutes behind and every downstream extract built on it is behind by the same amount plus its own schedule.
What the user sees
Numbers that are correct and old, presented identically to numbers that are correct and current. That is the actual defect: the product publishes a value without publishing its as-of time. Freshness was written in an SLO document rather than in the interface.
What stops it
- Publish freshness as data. An
as_oftimestamp on every row or partition and a watermark exposed in the product's own interface, so a consumer can decide. This is the change that turns a silent failure into a visible one and it costs one column. - Declare a peak tier in the contract. For example 60 minutes during announced events rather than 15, agreed in advance. A degradation nobody agreed to is an incident; the same degradation agreed in the contract is a mode.
- Shed the tail deliberately. Compute the top few thousand merchants by volume at 15 minutes and the remainder hourly. Merchant sales distributions are heavily skewed, so this typically recovers most of the cost while leaving the merchants who are watching unaffected.
- Make the aggregation incremental per window rather than a full recompute. Without this the system cannot self-heal, because catching up costs more than keeping up.
What would have to be true for it to self-heal
Only one thing: the cost of a window must not depend on how far behind you are. Incremental aggregation over a bounded window satisfies that. A recompute over a growing lookback does not, and it is the reason lag from a six-hour event can still be present the next morning.
When this is the wrong answer
If the consumers are finance and the product feeds monthly settlement, freshness degradation costs nothing and the engineering belongs elsewhere: on reconciliation and correctness. The as-of column is still worth having. The peak tier, the shedding and the incremental rewrite are not.