Semantic IDs for Recommendation
Why recommenders hash item identifiers into meaningless buckets, how quantising a content embedding into a short tuple of discrete codes replaces that hash with an address that carries similarity, and what the second training stage costs you.
A production ranking model at YouTube maps a corpus on the order of 100 million videos into an embedding table with on the order of 10 million buckets, using a random hash (Singh et al., 2024, Better Generalization with Semantic IDs, RecSys '24). The hash is random on purpose: it spreads load evenly across rows. It also guarantees that two videos about the same subject, by the same creator, uploaded an hour apart, land in unrelated rows and share no gradient. Every item has to earn its representation from its own impressions, which is fine for the head of the catalogue and hopeless for the tail.
A semantic ID replaces that hash with an address derived from what the item is. Take a content embedding from a frozen encoder, quantise it into a short tuple of discrete codes, and use the tuple as the identifier. Items that are similar in content now share a prefix, so they also share embedding rows, so they share statistical strength.
Residual quantisation, level by level
The standard construction is a residual-quantised variational autoencoder, borrowed from image tokenisation (Lee et al., 2022, Autoregressive Image Generation using Residual Quantization, CVPR 2022). Encode the item to a latent vector \(z\). Find the nearest entry \(c_1\) in the first codebook, record its index, and subtract:
Then quantise the residual \(r_1\) against a second codebook, subtract again, and repeat for \(L\) levels. The item's identifier is the tuple of indices \((i_1, i_2, \ldots, i_L)\).
Two properties follow from the recursion, and both matter more than the reconstruction error. The scheme is coarse to fine: the first code picks a broad region of content space, each later code refines it, so a shared prefix genuinely means "similar". And the address space is multiplicative while the parameter cost is additive. With \(L = 3\) levels of \(K = 256\) codes, TIGER addresses roughly 16.7 million distinct items using 768 codebook rows (Rajput et al., 2023, Recommender Systems with Generative Retrieval, NeurIPS 2023). A one-row-per-item table for the same catalogue would be four orders of magnitude larger.
Two quite different things to do with the result
The tuple can be a feature, dropped into an existing ranking model beside everything else. Singh and colleagues did this inside the YouTube ranker and found that treating the code sequence as text, and learning sub-piece units over it with SentencePiece, beat hand-crafted n-gram pieces and beat random hashing, with the gap widest on new and long-tail item slices and no loss on overall quality.
Or the tuple can be the generation target, which is what makes generative retrieval possible at all: a decoder emits \(i_1\), then \(i_2\), then \(i_3\), and the item falls out. That use is covered in generative retrieval for recommendation, and it is a much bigger commitment than the feature use.
Conflating the two is the most common mistake in this area. The feature version is a low-risk representation change that can be A/B tested against the existing hash. The generation version replaces the retrieval mechanism.
When it breaks
Collisions are guaranteed and must be handled explicitly. Nothing stops two items from quantising to the same tuple. The usual fix is to append an extra disambiguating token, which means the identifier length is no longer fixed by the codebook design and grows with catalogue density.
The tokeniser is a separate model with its own training run. It sits upstream of the recommender and is usually fitted against a frozen content encoder. That is two artefacts to version, two to monitor, and a coupling that only shows up when one of them changes: swap the encoder and every identifier in the catalogue changes, which is the failure mode described in item tokeniser drift and catalogue updates.
Codebook capacity saturates, and more encoder does not fix it. Scaling the content encoder from 77M to 11B parameters produced virtually no downstream gain, and enlarging the codebooks helped only up to a plateau (Zheng et al., 2025, Understanding Generative Recommendation with Semantic IDs from a Model-scaling View, arXiv:2509.25522). A few hundred discrete codes is a narrow channel, and it is the channel, not the encoder, that binds.
Content similarity is not behavioural similarity. Two films with near-identical synopses can have disjoint audiences. Semantic IDs import the encoder's notion of similarity wholesale, including the parts that do not predict engagement, which is why the strongest results come from systems that keep a collaborative signal alongside rather than replacing it.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- Singh et al., 2024, Better Generalization with Semantic IDs, RecSys '24 dl.acm.org
- Lee et al., 2022, Autoregressive Image Generation using Residual Quantization, CVPR 2022 openaccess.thecvf.com
- Rajput et al., 2023, Recommender Systems with Generative Retrieval, NeurIPS 2023 arxiv.org
- Zheng et al., 2025, Understanding Generative Recommendation with Semantic IDs from a Model-scaling View, arXiv:2509.25522 arxiv.org
6 flashcards for this concept
Click a card to reveal the answer.