Feature Freshness
How stale a feature can be before the model degrades, and the pipeline that follows.
4 to work through
-
advanced
ByteDance's Monolith paper (ORSUM at RecSys 2022) describes a collisionless embedding table for a recommender trained on the interaction stream, where new user and item ids arrive continuously so the key space has no ceiling. What bounds the memory, and what does each bound cost in model quality?
3 min answer -
advanced
ByteDance's Monolith paper (RecSys workshop 2022) describes training recommendation models on user feedback as it arrives rather than in nightly batches, and states that system reliability was deliberately traded for real-time learning. What does that trade actually look like in the pipeline, and when is a nightly batch the better engineering decision?
2 min answer -
advanced
Delivery ETAs are computed by a model using live traffic and restaurant load. The model's features are computed in batch overnight. What is wrong?
2 min answer -
advanced
How should feature freshness requirements be decided, and what does getting them wrong cost?
2 min answer
3 terms in this topic
Feature Freshness
How recent the data behind a model's input features is at the moment of inference, and the training-serving consistency problem it creates.
metricFeature Staleness
The age of the feature values a model scores against, and the divergence between how they are computed at training time and at inference time.
patternOnline Training Loop
Training a model continuously on interaction events as they arrive rather than on a scheduled batch, trading the ability to gate a model artefact bef…
Neighbouring topics
Streaming & Real-Time Data
General material on continuous processing of unbounded data.
Streaming vs Batch
The freshness requirement that actually justifies streaming, and the cost of assuming one.
Exactly-Once Semantics
What the phrase really means, where it holds, and the idempotent sink underneath it.
Stream Processing Frameworks
Flink, Kafka Streams, Spark Structured Streaming — state, checkpointing and recovery.
Windowing
Tumbling, sliding and session windows, and the aggregation each one answers.
Watermarks & Late Data
Deciding a window is complete when events can still arrive, and what to do when they do.
Stateful Stream Processing
Keyed state, state backends, checkpoint size, and the restore time that follows.
Stream-Table Duality
A changelog and a table as two views of the same thing, and materialising between them.
Kappa vs Lambda
One pipeline replayed versus two pipelines reconciled, and the maintenance each carries.
Streaming Schema Evolution
Changing an event's shape while a retained log still holds every older version of it.
Streaming Joins
Joining two unbounded streams, the buffering it needs, and the enrichment alternative.
Backfill & Reprocessing
Replaying history through changed logic without double-counting the live output.
CDC to Stream
Turning database changes into an event log, and how that differs from a domain event.
Real-Time Serving Layer
Where a low-latency read of a streaming aggregate actually lands.
Streaming SLOs
End-to-end latency, consumer lag and completeness as commitments rather than dashboards.
Partition Keys & Ordering
Ordering guaranteed only within a partition, and choosing the key that makes that enough.
Dead Letter Handling
The poison message that blocks a partition, and the queue nobody reads.
Streaming Cost
Always-on compute, retention and cross-zone traffic as the three bills that surprise.
Real-Time Analytical Stores
Druid, Pinot and ClickHouse — ingest-and-query engines for sub-second aggregation.