Model Architecture

The Generative Turn in Recommender Systems: Three Changes Wearing One Name

Kuaishou replaced its entire retrieve-rank cascade with one decoder and reported running it at 10.6 percent of the old system's operating expense. A controlled comparison of the same model family found near-zero recall on cold-start items. Both results are real, because 'generative recommendation' names three separable changes and only one of them is carrying the production wins.

Kuaishou reports that OneRec, a single encoder-decoder that reads a user's behaviour and emits a session of videos, replaced its cascaded recommender in two surfaces serving 400 million daily active users, lifted watch time by 1.6 percent, and ran at 10.6 percent of the previous pipeline's operating expense (OneRec Team, 2025, arXiv:2502.18965). That is not a benchmark delta. That is a production system whose bill fell by an order of magnitude.

Eight months earlier, a careful head-to-head against sequential dense retrieval found that generative retrieval, the technique at the heart of that architecture, "achieved near-zero performance" on cold-start items (Yang et al., 2024, Unifying Generative and Dense Retrieval for Sequential Recommendation, arXiv:2411.18814).

Both findings hold up. They are compatible because the phrase "generative recommendation" has been doing the work of three separate engineering changes, and the one that earns the production numbers is not the one that gives the field its name.

Why this matters: If you are deciding whether to rebuild a recommender around this, the single most valuable thing you can do is refuse the package deal. One of the three changes is a low-risk representation swap with clean evidence. One is a scaling story that has little to do with generation. One replaces your retrieval mechanism and brings a failure mode that aggregate recall does not show you.

TL;DR

  • "Generative recommendation" bundles three independent changes: content-derived item addresses (semantic IDs), decoding an identifier instead of scoring a catalogue (generative retrieval), and modelling raw action streams instead of engineered features (sequential transduction). They can be adopted separately and mostly should be.
  • Semantic IDs as a ranking feature are the safest win. Google reports they replace random video IDs in YouTube's ranker with better generalisation on new and long-tail slices and no loss overall (Singh et al., RecSys '24).
  • Generative retrieval's structural advantage is genuine: beam search over a code tree costs O(B · L · K) regardless of catalogue size. Its quality advantage is contested, and it collapses on new items.
  • The largest reported production gains come from the third change. Meta's HSTU hit a 12.4 percent online A/B improvement with 1.5 trillion parameters by reformulating ranking as next-action prediction, and its serving algorithm runs a model 285x more expensive in FLOPs at 1.50–2.48x higher throughput than the DLRM baseline.
  • Scaling laws hold for sequence models over action streams (98.3K to 0.8B parameters, power law, even data-constrained) but not for semantic-ID pipelines: scaling the content encoder from 77M to 11B parameters produced virtually no gain, because the identifier is the bottleneck.
  • Moving the index into the weights makes every catalogue insert a training problem. In the document-search version of this architecture, sequential indexing drops retrieval accuracy on previously indexed items from roughly 80 percent to below 20 percent.
  • The decisive adoption question is not quality. It is your catalogue churn rate divided by your retraining cadence.

At a Glance

flowchart LR
    A["Item content"] --> B["Frozen encoder"]
    B --> C["RQ-VAE tokeniser"]
    C --> D["Semantic ID tuple"]
    D --> E["Change 1: ID as feature"]
    D --> F["Change 2: ID as decode target"]
    G["Raw action stream"] --> H["Change 3: sequence model"]
    E --> I["Existing ranker, better tail"]
    F --> J["Index lives in weights"]
    H --> K["Compute now buys quality"]

    classDef blue fill:#1e40af,stroke:#3b82f6,stroke-width:1px,color:#fff
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
    classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
    class A,G blue
    class B,C,H purple
    class D,E,F teal
    class I,K teal
    class J amber

Three arrows leave the same tokeniser and go to genuinely different places. Most of the confusion in this literature comes from papers that take all three and report one number.

Before Anything Was Generated

The industrial recommender of 2019 was a funnel and a feature store. A retrieval stage pulled a few thousand candidates from millions, usually with two-tower retrieval and an approximate nearest-neighbour index. A ranking model scored those candidates from a very wide sparse feature vector, and the canonical form of that model put almost all its parameters in embedding lookup tables with a comparatively small dense network on top (Naumov et al., 2019, Deep Learning Recommendation Model, arXiv:1906.00091).

That design had a property nobody liked to say out loud: it did not respond to compute. Almost none of the parameters participate in a matrix multiply, so the system is bound by memory traffic and communication. Ten times the FLOPs bought very little. Progress came from feature engineering, and feature engineering does not have a scaling curve.

The sequence-modelling thread ran alongside it. SASRec made a user's history a transformer input in 2018 (Kang & McAuley, ICDM '18), BERT4Rec added bidirectional masked training the year after (Sun et al., CIKM '19), and both stayed academic-scale for years, partly because the protocol used to evaluate them made progress hard to read; see the progress illusion in recommender systems.

The unlock came from two places outside recommendation. Residual quantisation gave a way to turn a continuous vector into a short tuple of discrete codes, developed for image generation (Lee et al., CVPR 2022). And the differentiable search index showed that a single transformer could map queries straight to document identifiers, with all corpus information in the weights (Tay et al., 2022, NeurIPS 2022). Put those together and you get an item address that means something and a model that can emit it.

timeline
    title How the generative turn assembled
    2018 : SASRec makes user history a transformer sequence
    2019 : DLRM formalises sparse tables plus small dense network
         : BERT4Rec adds bidirectional masked training
    2022 : Differentiable Search Index maps queries to docids
         : RQ-VAE stacks residual codebooks for image tokens
    2023 : TIGER decodes semantic IDs for recommendation
         : DSI++ measures forgetting in a neural index
    2024 : HSTU reformulates ranking as sequential transduction
         : Semantic IDs enter YouTube ranking as features
         : LIGER finds near-zero cold-start generative recall
    2025 : OneRec collapses the Kuaishou cascade end to end
         : Netflix consolidates on a recommendation foundation model
         : Meta ships generative ads ranking with GEM
    2026 : Tree-decoding expressiveness limits formalised

[IMAGE: Two-panel schematic. Left panel "2019": a wide flat feature vector feeding a small MLP, with a huge embedding table drawn to scale beside it occupying most of the panel. Right panel "2024": a long horizontal token stream of interleaved item and action tokens feeding a deep stack, with the embedding table now a small strip. Caption: "The parameters did not shrink. They moved from lookup tables the accelerator cannot use into layers it can."]

Three Changes Wearing One Name

Change one: the address

A production ranker has to represent every item in the catalogue. The standard method hashes item identifiers into a fixed number of buckets: YouTube's ranker maps a corpus on the order of 100 million videos into on the order of 10 million buckets. The hash is random on purpose, because random spreads load evenly. It also guarantees that two near-identical items share no gradient.

A semantic ID derives the address from the content instead. Encode the item to a latent vector \(z\), then quantise it recursively. Find the nearest entry \(c^{(1)}_{i_1}\) in the first codebook, record the index, and take the residual:

\[r_1 = z - c^{(1)}_{i_1}, \qquad r_\ell = r_{\ell-1} - c^{(\ell)}_{i_\ell}\]

After \(L\) levels the identifier is the tuple \((i_1, \ldots, i_L)\). Two properties fall out of the recursion, and the reconstruction error is the less interesting one.

The scheme is coarse to fine. Level one picks a broad region of content space; each later codebook corrects what the previous level missed. A shared prefix therefore means real similarity, and a model that shares embedding rows across a prefix is sharing statistical strength between genuinely related items.

The address space is multiplicative while the cost is additive. With \(L = 3\) and \(K = 256\), TIGER addresses about 16.7 million items from 768 codebook rows (Rajput et al., 2023, NeurIPS 2023). One row per item for the same catalogue is four orders of magnitude more parameters, each learned alone.

Used as a feature, this is a contained change. Singh and colleagues dropped semantic IDs into YouTube's existing ranking model, found that treating the code sequence as text and learning sub-piece units over it with SentencePiece beat hand-crafted n-gram pieces, and reported better generalisation on new and long-tail item slices without sacrificing overall quality. Nothing about the serving architecture moved.

Change two: the mechanism

The second change is the one the word "generative" actually refers to, and it is much larger. Instead of scoring a catalogue, a sequence-to-sequence model emits the identifier: it reads the user's history as a flat stream of semantic IDs and decodes \(i_1\), then \(i_2\), then \(i_3\).

The retrieval cost argument is clean. Beam search of width \(B\) over \(L\) levels of \(K\)-way codebooks visits \(O(B \cdot L \cdot K)\) continuations, and none of those constants has anything to do with how many items you have. The ANN index goes away; what replaces it is decoder parameters that were going to exist anyway.

The decode has to be constrained. Most of the code tree addresses nothing, so an unconstrained decoder emits tuples for items that do not exist, and those occupy beam slots that would otherwise hold real candidates. Production systems mask each step's logits to the children present in a prefix trie over the live catalogue, which quietly reintroduces a catalogue-sized structure on the serving path, though a far cheaper one than a vector index.

TIGER reports Recall@5 of 0.0441 and NDCG@5 of 0.0309 on Amazon Beauty, with gains up to 29 percent in NDCG@5 over SASRec and 17.3 percent in Recall@5 over S³-Rec.

Change three: the substrate

The third change does not require semantic IDs or decoding at all, and it is where the largest production numbers come from. Meta's Hierarchical Sequential Transduction Unit throws away the heterogeneous feature vector and represents everything a user did as one time-ordered token stream, interleaving items with the actions taken on them. Ranking and retrieval both become sequential transduction (Zhai et al., 2024, Actions Speak Louder than Words, ICML 2024).

Three things travel together in that paper and separating them matters. The reformulation is the idea. The architecture is an attention unit built for long sequences and non-stationary vocabularies, running 5.3x to 15.2x faster than FlashAttention2-based transformers at length 8192. The serving algorithm, M-FALCON, scores candidates in microbatches under a mask that stops them attending to each other, so the attention over the user's history is computed once and amortised across the batch; at 1024 and 16384 candidates that yields 1.50x and 2.48x higher throughput than the DLRM baseline despite a model 285 times more expensive in FLOPs.

The headline outcome is a 1.5-trillion-parameter model and a 12.4 percent online A/B improvement, deployed on multiple surfaces of a platform with billions of users.

Why the three get conflated

Because the systems that report the biggest wins use all three, and because the word "generative" attaches naturally to any of them. But they have different risk profiles, different prerequisites and different evidence. Change one is testable next sprint. Change three is an infrastructure programme with a scaling curve behind it. Change two is the one that alters what your system can and cannot retrieve, and it is the one whose independent evidence is weakest.

[IMAGE: A decision tree for practitioners. Root: "What are you trying to fix?" Branches: "Tail and cold-item quality" leads to "Semantic IDs as a feature (low risk)"; "Model quality is flat despite more compute" leads to "Sequential transduction (infrastructure programme)"; "ANN index cost and complexity" leads to "Generative retrieval (check churn rate first)". Caption: "The three changes answer three different questions. Adopting all three because one of them fits is the common error."]

Seeing It in Motion

A hybrid retrieval request, which is what most production systems that use any of this actually run:

sequenceDiagram
    participant U as User request
    participant S as Sequence encoder
    participant D as Semantic ID decoder
    participant T as Prefix trie
    participant X as Cold-item index
    participant R as Dense reranker
    U->>S: Recent action history
    S->>D: User state vector
    D->>T: Beam step, request valid children
    T-->>D: Allowed code set
    Note over D,T: Repeat once per codebook level
    D->>R: Generated candidate IDs
    X->>R: Recent and cold items injected
    R-->>U: Final ranked list

The trie round trips are the part that surprises people. The decode is cheap in FLOPs and expensive in dependencies: each level needs the previous level's choice before it can start.

The structural comparison between a cascade and an end-to-end generative recommender:

flowchart TB
    subgraph Cascade["Staged cascade"]
        C1["Retrieve thousands"] --> C2["Pre-rank to hundreds"]
        C2 --> C3["Rank with heavy model"]
        C3 --> C4["Policy and diversity rules"]
    end
    subgraph EndToEnd["End-to-end generative"]
        E1["Encode user history"] --> E2["Decode session directly"]
        E2 --> E3["Reward-aligned output"]
    end
    C4 --> O1["Calibrated, auditable, slow to change"]
    E3 --> O2["Efficient, opaque, retrain to retune"]

    classDef slate fill:#334155,stroke:#64748b,stroke-width:1px,color:#e2e8f0
    classDef purple fill:#6d28d9,stroke:#a78bfa,stroke-width:1px,color:#fff
    classDef teal fill:#0e7490,stroke:#22d3ee,stroke-width:1px,color:#fff
    classDef amber fill:#b45309,stroke:#fbbf24,stroke-width:1px,color:#fff
    class C1,C2,C3,C4 slate
    class E1,E2,E3 purple
    class O1 teal
    class O2 amber

Every arrow inside the cascade is also a place where a filter runs. That is the property being traded away, and it does not appear in any efficiency table.

Watch It Run

Animated diagram showing an item embedding quantised through three residual codebooks into a semantic ID, then a user history decoded level by level into candidate identifiers, with a feedback loop from cold-item injection back into the reranker.
Solid animated edges carry the offline tokenisation path (content to embedding to residual codebooks to semantic ID) and the online decode path (history to beam step to candidates). The animated self-loop on the beam-step node is the per-level decode repeating once per codebook. The amber feedback edge is cold-item injection, the path that exists precisely because the decoder cannot reach new items on its own. The static Mermaid figures above show the same structure if the animation is absent.

By the Numbers

System Change(s) used Reported result Compared against
TIGER Semantic ID + generative retrieval Recall@5 0.0441, NDCG@5 0.0309 (Amazon Beauty); up to +29% NDCG@5 SASRec; S³-Rec (+17.3% Recall@5)
HSTU (public impl.) Sequential transduction +23.1% HR@10, +29.4% NDCG@10 (MovieLens-20M); +56.7% / +60.7% (Amazon Books) SASRec
HSTU (production) Sequential transduction +12.4% online A/B metrics at 1.5T parameters Prior production ranker
M-FALCON Serving algorithm 1.50x / 2.48x throughput at 1024 / 16384 candidates, on a 285x-FLOPs model DLRM serving baseline
OneRec All three, end to end 23.7% / 28.8% MFU train / inference; 10.6% of prior OPEX; +1.6% watch time at 400M DAU Kuaishou's own cascade
Meta GEM Sequential transduction (ranking layer) 4x more efficient per unit data and compute; 23x effective training FLOPs on 16x GPUs Prior ads ranking models
Sequential scaling study Sequential transduction Power law from 98.3K to 0.8B parameters; 0.8B predicted from sub-100M models Within-family scaling
Semantic ID scaling Semantic ID pipeline Encoder 77M to 11B: virtually no gain; codebook growth plateaus Within-family scaling
LIGER Generative vs dense, controlled Generative retrieval near-zero on cold-start items Sequential dense retrieval
DSI++ Neural index, continual update Indexing accuracy ~80% to below 20% under sequential indexing; +21.1% Hits@10 with replay, 6x fewer updates Retraining from scratch
Latte Tree-decode fix +3.45% NDCG@10 average from latent-token decoding Standard semantic-ID decoding

Sources: TIGER (Rajput et al., 2023); HSTU and M-FALCON (Zhai et al., 2024, public numbers from meta-recsys/generative-recommenders); OneRec (arXiv:2502.18965, arXiv:2506.13695); Meta GEM (Meta Engineering, 2025); scaling law (Zhang et al., 2024); semantic ID scaling (arXiv:2509.25522); LIGER (Yang et al., 2024); DSI++ (Mehta et al., 2023); Latte (arXiv:2605.06331).

Every production figure here is vendor-reported against that vendor's own previous system. None of them is an independent measurement, and the online A/B numbers in particular cannot be reproduced by anyone outside the company. Treat them as existence proofs at scale, not as margins.

A Concrete Example

Take a catalogue of 2,000,000 items and a tokeniser with \(L = 3\) levels of \(K = 256\) codes.

Step 1: tokenise one item. Its content encoder gives a latent vector, shown here in four dimensions so the arithmetic is readable:

z  = [ 0.82, -0.31,  0.45,  0.10 ]    ||z||  = 0.991

Nearest entry in codebook 1 is index 17, equal to [0.80, -0.30, 0.40, 0.00]. Subtract:

r1 = [ 0.02, -0.01,  0.05,  0.10 ]    ||r1|| = 0.114

Level one has captured 88.5 percent of the vector's norm in a single integer. Nearest entry in codebook 2 is index 203, equal to [0.00, 0.00, 0.06, 0.09]:

r2 = [ 0.02, -0.01, -0.01,  0.01 ]    ||r2|| = 0.027

Level two takes the residual to 2.7 percent of the original norm. Codebook 3 index 88 absorbs most of what is left. The semantic ID is (17, 203, 88), three integers standing in for a dense vector, and the first of them alone already says roughly what the item is about.

Step 2: count the address space. \(256^3 = 16{,}777{,}216\) addresses for 2,000,000 items, so 11.9 percent occupancy. Under a uniform-random assumption the expected number of occupied addresses is

\[m\left(1 - e^{-n/m}\right) = 16{,}777{,}216 \times \left(1 - e^{-0.1192}\right) \approx 1{,}885{,}500\]

which leaves roughly 114,500 items, about 5.7 percent, sharing an address with something else. Those need a disambiguation token. Real semantic IDs cluster rather than spreading uniformly, so 5.7 percent is a floor, not an estimate of the collision rate you will see.

Step 3: decode for a user. With beam width 20, the first step scores 256 candidates and keeps 20. The second scores \(20 \times 256 = 5{,}120\). The third scores another 5,120, of which only the ones present in the trie are legal: with 2,000,000 leaves spread over roughly 65,536 level-two prefixes, a typical prefix has about 30 real children out of 256, so the mask discards roughly 88 percent of that final step's candidates. Total continuations considered: about 10,500, against 2,000,000 items for exhaustive scoring. That is the catalogue-independence argument, in numbers.

It is also three dependent forward passes, one per level, inside a retrieval budget usually measured in tens of milliseconds.

Step 4: now add a new item. A film uploaded after the last training run tokenises fine, to say (17, 203, 91). The prefix (17, 203) is well-travelled, so the first two decode steps are no obstacle. The third is. The model's distribution over third-level codes given that prefix was fitted on codes it saw, and 91 was not one of them, so its logit sits near the initialisation floor. Suppose the 20th-best cumulative log-probability in the beam is \(-6.2\) and the new item's path scores \(-11.4\): it never enters the beam, at any beam width you can afford.

These log-probabilities are illustrative. The measured result they illustrate is not: in a controlled comparison the minimum generation probability in the beam exceeded the probability assigned to ground-truth cold-start items, producing near-zero recall for them.

Step 5: what a hybrid does about it. Generate the 20 candidates, pull recent and cold items from a small dense index, and re-rank the union with a dense scorer. You keep the storage win and recover the cold-start recall, and you are now running two retrieval systems instead of one.

[IMAGE: Horizontal stacked bar showing residual norm after each quantisation level for the worked example: 100% before, 11.5% after level 1, 2.7% after level 2, near 0% after level 3, with each bar annotated with the chosen codebook index. Caption: "Coarse to fine: the first integer does most of the work, which is why a shared prefix means something."]

[IMAGE: Beam search trellis over three codebook levels, with 256 nodes at level 1 narrowing to a beam of 20, expanding to 5,120 at level 2, and at level 3 showing greyed-out nodes that the prefix trie masks away. One node at level 3, coloured amber and unreachable, is labelled "new item, code 91, never seen in training". Caption: "The constrained decode is what makes the beam efficient. It is also what makes a new item invisible."]

Where It Breaks

Cold-start failure is structural, not a tuning problem

The decoder's conditional probabilities are fitted to identifier paths it walked. For a genuinely new item, the correct path scores below the beam's floor, so the item is not ranked poorly, it is unreachable. This is the precise inverse of the claim semantic IDs are usually sold on: the codes generalise to unseen items as features, because they are shared, but a decoder does not inherit that generalisation, because what it learned is a distribution over walked paths.

Tree structure correlates scores that should be independent

Items sharing a prefix share every factor in the probability chain up to their divergence point, so their scores move together whatever the user prefers. This is not a training artefact but an expressiveness limit: there exist simple user-item preference patterns and item-item similarity relations that plain collaborative filtering represents easily and an autoregressive semantic-ID decoder provably cannot. Injecting a latent token before each identifier, which splits the single decode tree into several latent-conditioned trees, recovers an average of 3.45 percent NDCG@10 (arXiv:2605.06331). The size of that fix tells you what the single tree was costing.

The scaling story does not transfer between the three changes

Sequence models over raw action streams show a power law from 98.3K to 0.8B parameters, holding even under severe data constraints, with the 0.8B model's performance predicted from models more than 100x smaller (Zhang et al., 2024, RecSys '24). Semantic-ID pipelines do not: performance saturates as each component is enlarged, the identifier's limited capacity is the named bottleneck, and pushing the content encoder from 77M to 11B parameters produced virtually no gain (arXiv:2509.25522). "Recommendation scales now" is true of one of the three changes.

Catalogue updates become a continual-learning problem

An ANN index accepts an insert; a decoder has to learn the path. In the document-search version of the architecture, indexing new batches sequentially drove retrieval accuracy on previously indexed documents from around 80 percent to below 20 percent, and the fix that worked was replay: sampling pseudo-queries for already-indexed documents and mixing them into the update stream, for +21.1 percent average Hits@10 and six times fewer model updates than retraining (Mehta et al., 2023). Recommendation catalogues turn over faster than document corpora.

A second, sharper version of the same problem: upgrade the content encoder and every identifier changes at once. There is no partial migration, because the new identifiers are a different language.

Deleting the cascade deletes its seams

The stage boundaries in a funnel are the only places where something other than a learned objective intervenes. Age gating, licensing windows, geographic restrictions and advertiser exclusions run as filters between stages; in an end-to-end model they have to become decode masks or reward penalties, both harder to audit than a filter that either ran or did not. Calibration goes too: an auction needs a probability, and a sequence likelihood is monotone in preference without being the probability of any event. So does diagnosis, since "the item was never shown" stops resolving to a stage and starts resolving to the weights.

The cadence cost is the one teams feel first. Re-weighting watch time against diversity is a config change in a cascade and a reward-shaping change plus a training run end to end, which turns a same-day experiment into a same-quarter one.

Reward alignment imports reward hacking

Once the generated session is optimised against a learned reward model, the recommender inherits reward over-optimisation in full: it will find the region of session space the reward model scores generously and users do not. A cascade's re-ranking policy sat downstream as a partial check on exactly this. Collapsing the cascade removes it.

The benchmarks are the same ones that misled the field before

Most offline results here are reported on Amazon Reviews and MovieLens splits, with sampled metrics and baselines of uneven tuning. That protocol has a documented history of producing gains that do not survive careful re-evaluation, covered in the progress illusion in recommender systems. A 29 percent NDCG@5 improvement on Amazon Beauty is a reason to run an experiment, not a reason to plan a migration.

Alternative Designs

Design How it works Key advantage Key limitation Best when
Two-tower plus ANN Embed user and items, nearest-neighbour search Mature, trivial inserts, one round trip Index memory and build pipeline; no cross-features Default, and the baseline anything else must beat
Semantic IDs as features Content codes replace or augment hashed IDs in the existing ranker Tail and cold-item gains with no architectural change Tokeniser is a second versioned artefact You have a long tail and a working cascade
Generative retrieval Decoder emits the identifier under a trie mask Retrieval cost independent of catalogue size Near-zero cold-start recall; retrain to insert Catalogue is large, slow-changing, and index cost hurts
Hybrid generate-then-rerank Generate candidates, inject cold items, dense re-rank Recovers cold start while keeping small storage Two retrieval systems to operate You want the generative path in production now
Sequential transduction One action-stream sequence model for ranking Compute converts to quality; large reported online gains Long-history cost; storage and bandwidth on critical path Feature engineering has plateaued and you have GPUs
End-to-end generative Single model decodes the output session Large efficiency and OPEX wins reported Loses filters, calibration, per-stage diagnosis Single-surface, single-objective products
LLM as recommender A language model scores or generates items directly World knowledge; better scaling than SID pipelines Latency and cost; collaborative signal is contested Cold-start-heavy or explanation-heavy surfaces

The row that deserves the most attention is the hybrid, because it is the one the evidence currently supports and the one least likely to appear in a keynote.

How It Is Used in Practice

Three deployment shapes are visible in public reporting, and they are not converging.

Feature-level adoption. Google put semantic IDs inside the existing YouTube ranking model, as a replacement representation for video IDs, and reported generalisation gains on new and long-tail slices with no loss overall. Nothing about retrieval or serving changed. This is the shape most organisations should copy first.

Layer replacement. Meta rebuilt the layers inside its advertising stack while keeping the structure: a retrieval engine feeding a generative ranking model, rather than one network emitting the final list. The reported efficiency claim is that the ranking model converts data and compute into ad performance about four times more efficiently than its predecessors, on a training stack delivering 23 times the effective FLOPs using 16 times the GPUs (Meta Engineering, 2025). Netflix took a related path from the other direction, consolidating many specialised models into one foundation model that downstream applications fine-tune from, which centralises learning while leaving the serving pipeline intact (Hsiao, Feng & Lamkhede, 2025, Netflix TechBlog).

Full collapse. Kuaishou is the clearest case of replacing the cascade outright, and its argument is as much about hardware as quality: in its own analysis, over half of serving resources in the conventional stack went to communication and storage rather than high-precision computation. A later report describes expansion to roughly 25 percent of total QPS (OneRec Team, 2025, arXiv:2506.13695), which is a meaningful qualifier on "replaced".

The operational question that decides which shape fits is not in any of these papers. It is the ratio of catalogue churn to retraining cadence. If a material share of impressions lands on items younger than your retraining interval, a purely generative retriever cannot serve that share at all. Short video, news, marketplaces and live commerce sit on the wrong side of that line; a film catalogue does not, which is a large part of why the deployment shapes differ by industry rather than by conviction.

[IMAGE: Scatter plot with catalogue churn rate on the x-axis (log scale, items added per day as a fraction of catalogue) and retraining cadence on the y-axis (hours between retrains). A diagonal feasibility boundary separates "hybrid or dense required" from "pure generative viable", with labelled points for a film catalogue, a music catalogue, a marketplace, and a short-video feed. Caption: "The adoption question is an operational ratio, not a quality comparison."]

[IMAGE: Three-column table rendered as a graphic, one column per deployment shape (feature-level, layer replacement, full collapse), with rows for "what changes", "what you keep", "who reported it", "reversibility". Caption: "Reversibility is the row that should be read first."]

Insights Worth Remembering

  1. "Generative recommendation" is three changes, and only one of them is generative. Content-derived addresses, identifier decoding, and action-stream sequence modelling are independent. The production wins are concentrated in the third, which does not require the other two.

  2. Semantic IDs generalise; semantic ID decoders do not. Sharing codes across items is what gives the representation its tail behaviour. Decoding is fitted to observed paths and throws that generalisation away, which is why cold-start recall collapses even though the addresses are content-derived.

  3. The identifier is the bottleneck, not the encoder. A few hundred codes per level is a narrow channel. Scaling the content encoder from 77M to 11B parameters bought virtually nothing, so if semantic-ID quality is your ceiling, a bigger embedding model is the wrong purchase.

  4. In candidate scoring, the per-user work is the amortisable part. M-FALCON's 285x-FLOPs model runs faster than the baseline purely because history attention does not depend on which candidate is scored. That observation generalises well beyond generative recommenders.

  5. Retrieval cost independence is bought with serial latency. A four-level decode is four dependent forward passes where an ANN lookup is one round trip. FLOP counts flatter this architecture; wall-clock does not.

  6. Moving the index into the weights converts an insert into a training run. Everything downstream of that, the cold-start collapse, the forgetting under sequential indexing, the tokeniser version coupling, is a consequence of that one decision.

  7. The seams of a cascade are load-bearing. Filters, calibration and per-stage diagnosis all live on stage boundaries. Efficiency tables never price them, and they are the part teams miss most after a collapse.

  8. Vendor figures here are existence proofs, not margins. Every production number in this area is measured against that company's own legacy stack. They show the architecture can work at scale; they say nothing about its advantage over a well-tuned modern pipeline.

Open Questions

Does generative retrieval have a cold-start fix that is not a hybrid? Measured: pure generative retrieval is near-zero on cold items, and generate-then-rerank with explicit cold injection recovers it. Unknown: whether any purely generative scheme reaches new items without a dense path beside it, since every credible published fix so far reintroduces one.

How much of the production gain is generation and how much is scale? The systems reporting the largest wins changed the representation, the mechanism and the substrate at once, against legacy baselines. No public ablation isolates the decoding change at production scale, and it is entirely possible that decoding contributes little beyond the efficiency of not running an index.

Can semantic IDs carry more without losing the properties that make them useful? The capacity bottleneck is measured. What is not settled is whether longer identifiers, larger codebooks or learned tokenisers can widen the channel without destroying the coarse-to-fine prefix semantics that make sharing work in the first place.

What replaces calibration in an end-to-end recommender? For organic feeds this may not matter. For anything with an auction, a budget or a contractual delivery guarantee, a sequence likelihood is not a substitute for a calibrated probability, and there is no published account of how a collapsed system rebuilds one.

Does the LLM-as-recommender path dominate the semantic-ID path? One study reports up to 20 percent improvement over the best achievable semantic-ID performance through scaling, and challenges the belief that language models cannot capture collaborative signal (arXiv:2509.25522). That is a single line of evidence on academic data, at latencies no large feed can currently afford, and it is the most interesting open question in the area.

Sources and Further Reading

  1. Rajput, S., et al. (2023). "Recommender Systems with Generative Retrieval." NeurIPS 2023. arXiv:2305.05065
  2. Zhai, J., et al. (2024). "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations." ICML 2024, PMLR 235. arXiv:2402.17152
  3. Tay, Y., Tran, V. Q., Dehghani, M., et al. (2022). "Transformer Memory as a Differentiable Search Index." NeurIPS 2022. arXiv:2202.06991
  4. Singh, A., et al. (2024). "Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations." RecSys '24. doi:10.1145/3640457.3688190, preprint arXiv:2306.08121
  5. Yang, L., et al. (2024). "Unifying Generative and Dense Retrieval for Sequential Recommendation." arXiv:2411.18814
  6. Mehta, S. V., et al. (2023). "DSI++: Updating Transformer Memory with New Documents." EMNLP 2023. ACL Anthology
  7. Lee, D., Kim, C., Kim, S., Cho, M., & Han, W.-S. (2022). "Autoregressive Image Generation using Residual Quantization." CVPR 2022. CVF Open Access
  8. Zhang, G., et al. (2024). "Scaling Law of Large Sequential Recommendation Models." RecSys '24. arXiv:2311.11351
  9. OneRec Team (2025). "OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment." arXiv:2502.18965
  10. OneRec Team (2025). "OneRec Technical Report." arXiv:2506.13695
  11. Zheng, J., et al. (2025). "Understanding Generative Recommendation with Semantic IDs from a Model-scaling View." arXiv:2509.25522
  12. Ju, C. M., Collins, L., Neves, L., Kumar, B., Wang, L. Y., Zhao, T., & Shah, N. (2025). "Generative Recommendation with Semantic IDs: A Practitioner's Handbook." arXiv:2507.22224; code at snap-research/GRID
  13. "Expressiveness Limits of Autoregressive Semantic ID Generation in Generative Recommendation" (2026). arXiv:2605.06331
  14. Naumov, M., et al. (2019). "Deep Learning Recommendation Model for Personalization and Recommendation Systems." arXiv:1906.00091
  15. Kang, W.-C., & McAuley, J. (2018). "Self-Attentive Sequential Recommendation." ICDM '18. arXiv:1808.09781
  16. Sun, F., et al. (2019). "BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer." CIKM '19. arXiv:1904.06690
  17. Meta Engineering (2025). "Meta's Generative Ads Model (GEM): The Central Brain Accelerating Ads Recommendation AI Innovation." engineering.fb.com
  18. Hsiao, K.-J., Feng, Y., & Lamkhede, S. (2025). "Foundation Model for Personalized Recommendation." Netflix TechBlog. netflixtechblog.com
  19. Meta (2024). "Generative Recommenders" reference implementation. github.com/meta-recsys/generative-recommenders

Free to read, no ads, no sign-up. If it was useful you can buy me a coffee.