Recommender Systems advanced 7 min read 7 flashcards

Item Tokeniser Drift and Catalogue Updates

Why moving the retrieval index into model weights turns every new item and every content-encoder upgrade into a continual-learning problem, and how much retrieval quality forgetting actually destroys.

An approximate nearest-neighbour index accepts an insert. Embed the new item, add the vector, and it is retrievable in milliseconds. That operation has no equivalent in a system whose index lives in decoder weights, and the gap is the single most under-discussed cost of generative retrieval.

The measurement comes from document search, where the same architecture was tried first. When a differentiable search index is updated by indexing new document batches in sequence, its ability to retrieve previously indexed documents collapses: indexing accuracy falls from around 80 percent to below 20 percent (Mehta et al., 2023, DSI++: Updating Transformer Memory with New Documents, EMNLP 2023). This is ordinary catastrophic forgetting, and it applies with more force to recommendation, where the catalogue turns over faster than any document corpus.

Two drifts, with different blast radii

Item churn is continuous and local. New items arrive, and the model has never decoded their identifiers. If they are addressable at all it is by luck, because the decoder's conditional probabilities are fitted to paths it has walked; this is the same mechanism behind the near-zero cold-start recall measured for generative retrieval (Yang et al., 2024, arXiv:2411.18814).

Tokeniser change is discrete and global. Upgrade the content encoder, refit the RQ-VAE, or resize a codebook, and every identifier in the catalogue changes at once. The decoder's learned vocabulary embeddings, and the transition structure it learned over the code tree, are both invalidated. There is no partial migration: the new identifiers are a different language, so you retrain, and until that retraining lands the two artefacts cannot be mixed.

This is why the tokeniser deserves to be treated as a versioned interface rather than a preprocessing step. A semantic ID is a contract between the tokeniser and the recommender, with all the compatibility questions that implies; the discipline in schema evolution and compatibility transfers directly.

What actually helps

Replay, not just more training. DSI++ attacks forgetting on two fronts: optimising for flatter minima so memorisation is more stable, and a generative memory that samples pseudo-queries for already-indexed documents and mixes them into the continual-indexing stream. Together they improve average Hits@10 by 21.1 percent over competitive baselines and need six times fewer model updates than retraining from scratch for five sequential corpora.

Reserve address space before you need it. Codebook capacity is fixed at tokeniser training time. Fitting the codebooks tightly to today's catalogue means tomorrow's items quantise into crowded regions and collide, so the disambiguation token does more work and the prefix carries less meaning.

Keep a path that does not need training. The hybrid designs exist partly for this reason: generate candidates, explicitly inject recent and cold items from a dense index, and re-rank the union. The dense side absorbs catalogue churn at index-insert cost while the generative side handles the head.

When it breaks

When churn outruns the retraining cadence. This is the decisive quantity and it is rarely stated. If a meaningful share of impressions falls on items younger than your retraining interval, a purely generative retriever cannot serve that share at all, and no amount of quality work on the head recovers it. News, short video, marketplaces and live commerce all sit on the wrong side of this line; a film catalogue does not.

When the encoder is someone else's. Content embeddings often come from a third-party or centrally-owned model that is upgraded on its own schedule. An upstream version bump becomes a forced retraining of every downstream recommender, and the coupling is usually discovered the first time it happens.

When forgetting is silent. Nothing errors. The model still emits valid identifiers at the same rate; they are simply the wrong ones, concentrated on the older parts of the catalogue. Aggregate recall moves a little, tail recall moves a lot, and without a per-cohort slice by item age the regression is invisible until a content partner complains.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Mehta et al., 2023, DSI++: Updating Transformer Memory with New Documents, EMNLP 2023 aclanthology.org
  2. Yang et al., 2024, arXiv:2411.18814 arxiv.org
Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track