Sequential Transduction for Recommendation
Why deep CTR models stopped improving when given more compute, how reformulating ranking as next-action prediction over one unified event stream restored a scaling curve, and where that curve flattens.
For roughly seven years the industrial recommender stack got better by adding features, not compute. Deep CTR models in the DLRM lineage combine enormous sparse embedding tables with a comparatively small dense network (Naumov et al., 2019, Deep Learning Recommendation Model, arXiv:1906.00091). Almost all the parameters sit in lookup tables; almost none participate in a matrix multiply. Give that architecture ten times the FLOPs and it does not get appreciably better, because the bottleneck was never arithmetic.
Sequential transduction is the reformulation that changes what the compute is spent on. Throw away the heterogeneous feature vector. Represent everything a user did as one time-ordered stream of tokens, interleaving items with the actions taken on them, and train the model to predict the next token. Ranking and retrieval both become instances of transducing one sequence into another.
What HSTU actually changed
Meta's Hierarchical Sequential Transduction Unit is the load-bearing example (Zhai et al., 2024, Actions Speak Louder than Words, ICML 2024). Three things travel together in that paper and they are worth separating, because only the first is the idea.
The reformulation unifies ranking and retrieval features into a single sequence, so a model that was previously reading a wide flat feature vector now reads a long history with position and time structure in it. The architecture replaces standard attention with a unit designed for the recommendation regime, where vocabularies are non-stationary and sequences are long: HSTU runs 5.3x to 15.2x faster than FlashAttention2-based transformers on sequences of length 8192. The serving algorithm, M-FALCON, scores candidates in microbatches with a mask that stops candidates attending to each other, so the user-history attention is computed once and amortised across all of them; scoring 1024 and 16384 candidates gives 1.50x and 2.48x higher throughput than the DLRM baseline despite a model 285 times more expensive in FLOPs.
The reported outcomes: a 1.5-trillion-parameter model, a 12.4 percent improvement in online A/B metrics, and deployment across multiple surfaces of a platform with billions of users. On public data, the released implementation reports HR@10 and NDCG@10 gains over SASRec of 15.5 and 18.1 percent on MovieLens-1M, 23.1 and 29.4 percent on MovieLens-20M, and 56.7 and 60.7 percent on Amazon Books (meta-recsys/generative-recommenders).
Does recommendation have a scaling law
Partly, and with caveats the headline invites you to skip. A systematic study over decoder-only sequential recommenders from 98.3K to 0.8B parameters found a power-law relationship that held across that range even under severe data constraints, and used sub-100M models to predict the 0.8B model's performance (Zhang et al., 2024, Scaling Law of Large Sequential Recommendation Models, RecSys '24). The same work reports that training instability grows with scale, which is the practical tax.
Meta's advertising stack reports the same directional finding from production: the Generative Ads Model is described as four times more efficient at converting data and compute into ad performance than the ranking models it replaced, on a training stack delivering 23 times more effective FLOPs using 16 times more GPUs (Meta, 2025, Meta's Generative Ads Model (GEM)). These are vendor-reported figures on an internal baseline, not an independent benchmark.
When it breaks
Quadratic attention meets unbounded histories. A heavy user has years of events. Truncating the sequence discards the long-range signal that motivated the reformulation; keeping it puts you on an \(O(n^2)\) curve at retrieval latency. Every production system here makes a truncation or sparsification choice, and that choice, not the architecture, usually sets the quality ceiling.
The scaling curve is architecture-specific, not universal. Generative recommenders built on semantic IDs saturate quickly as components are enlarged, with the identifier's limited capacity identified as a fundamental bottleneck (arXiv:2509.25522). "Recommendation scales" is true of sequence models over raw action streams; it is not yet true of every system that calls itself generative.
Storage and bandwidth move to the critical path. Serving a long per-user history means fetching a long per-user history. In Kuaishou's analysis of its own conventional stack, more than half of serving resources went to communication and storage rather than high-precision computation (OneRec Team, 2025, arXiv:2506.13695), and sequence models make that ratio worse before hardware makes it better.
Non-stationary vocabularies. Language models are trained on a fixed vocabulary. A catalogue turns over continuously, so the embedding table is a moving target and the token distribution shifts under the model while it trains.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Naumov et al., 2019, Deep Learning Recommendation Model, arXiv:1906.00091 arxiv.org
- Zhai et al., 2024, Actions Speak Louder than Words, ICML 2024 arxiv.org
- meta-recsys/generative-recommenders github.com
- Zhang et al., 2024, Scaling Law of Large Sequential Recommendation Models, RecSys '24 arxiv.org
- Meta, 2025, Meta's Generative Ads Model (GEM) engineering.fb.com
- arXiv:2509.25522 arxiv.org
- OneRec Team, 2025, arXiv:2506.13695 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.