Collapsing the Retrieval and Ranking Cascade
What a single end-to-end generative recommender buys in efficiency when it replaces the retrieve-prerank-rank pipeline, and what you give up by deleting the stage boundaries that business logic, calibration and debugging all live on.
Every large recommender is a funnel: retrieve a few thousand candidates from millions, pre-rank down to hundreds with a cheap model, rank those with an expensive one, then apply a re-ranking policy for diversity, freshness and business rules. Multi-stage ranking cascades exist for one reason, which is that you cannot afford the expensive model over the whole catalogue.
The end-to-end generative argument is that the funnel is an artefact of that constraint rather than a property of the problem, and that a single model trained to generate the output list directly will beat a pipeline of separately-optimised stages whose objectives do not agree.
The evidence for collapsing it
Kuaishou's OneRec is the clearest production case. It replaces the cascade with one encoder-decoder that reads the user's behaviour sequence and decodes a session of videos, uses a sparse mixture of experts to add capacity without proportional FLOPs, and aligns the output with an iterative preference procedure over a reward model (OneRec Team, 2025, OneRec: Unifying Retrieve and Rank with Generative Recommender, arXiv:2502.18965).
The reported numbers are about hardware efficiency as much as quality: 23.7 percent Model FLOPs Utilisation in training and 28.8 percent in inference, against a conventional pipeline where over half of serving resources go to communication and storage rather than computation, and total operating expense at 10.6 percent of the previous system's, with a 1.6 percent watch-time gain in deployment on scenarios serving 400 million daily active users. A later report describes expansion to roughly 25 percent of total QPS and, in the local-life-services scenario, GMV growth of 21.01 percent (OneRec Team, 2025, OneRec Technical Report, arXiv:2506.13695).
The efficiency argument deserves to be taken seriously on its own terms. A cascade of small models spends its budget moving embeddings around; a dense decoder spends it on matrix multiplies that accelerators are built for. That is a structural difference, not a tuning difference.
What the stage boundaries were carrying
The seams of a cascade are not only compute checkpoints. They are the only places where anything other than a learned objective can intervene.
Policy and compliance. Age gating, licensing windows, geographic restrictions, advertiser exclusions and "do not show this again" all execute as filters between stages. A decoder that emits a list has no such seam; the constraints have to become masks on the decode or penalties in the reward, and both are harder to audit than a filter that either ran or did not.
Calibration. An auction needs a probability. A ranking model trained with a logistic objective produces something that can be calibrated to a real click rate; a decoder produces a sequence likelihood, which is monotone in preference but is not a probability of any event. Anywhere a bid, a budget or an SLA depends on the number, calibration has to be rebuilt on top.
Diagnosis. In a funnel, "the item was never shown" resolves to a specific stage that dropped it. In an end-to-end model, it resolves to the weights.
This is why the industry has not moved uniformly. Meta's advertising stack kept its structure and replaced the layers inside it, pairing a retrieval engine with a generative ranking model rather than emitting the final list from one network (Meta, 2025). Netflix consolidated many specialised models into one foundation model that downstream applications fine-tune from, which centralises learning while leaving the serving pipeline in place (Hsiao, Feng & Lamkhede, 2025, Foundation Model for Personalized Recommendation, Netflix TechBlog).
When it breaks
Multi-objective trade-offs lose their dial. A cascade lets you weight watch time against diversity against creator equity at the re-ranking stage, and change the weights without retraining. Folded into a single generative objective, every such trade-off becomes a reward-shaping change and a training run, which converts a same-day experiment into a same-quarter one.
Reward hacking arrives with preference alignment. Once the output is aligned against a learned reward model, the recommender inherits the failure modes of reward over-optimisation: it will find the region of session space the reward model scores generously and the users do not.
Blast radius. One model serving retrieval and ranking is one bad checkpoint away from a total outage of the surface. A cascade degrades: lose the ranker and you serve retrieval order.
The comparison is rarely clean. Reported end-to-end wins are measured against that company's own legacy cascade, which had years of accumulated patches and a different compute budget. Treat them as existence proofs that the architecture can work at scale, not as a measured margin over a well-tuned modern pipeline.
References and further reading
Every source this page cites, in the order it cites them. All of them open in a new tab.
- OneRec Team, 2025, OneRec: Unifying Retrieve and Rank with Generative Recommender, arXiv:2502.18965 arxiv.org
- OneRec Team, 2025, OneRec Technical Report, arXiv:2506.13695 arxiv.org
- Meta, 2025 engineering.fb.com
- Hsiao, Feng & Lamkhede, 2025, Foundation Model for Personalized Recommendation, Netflix TechBlog netflixtechblog.com
7 flashcards for this concept
Click a card to reveal the answer.