Full Stage Learning to Rank: A Unified Framework for Multi-Stage Systems
Kai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang, Na Mou, Yanan Niu, Yang Song, Hongning Wang, Kun Gai
Abstract
The Probability Ranking Principle (PRP) has been considered as the foundational standard in the design of information retrieval (IR) systems. The principle requires an IR module's returned list of results to be ranked with respect to the underlying user interests, so as to maximize the results' utility. Nevertheless, we point out that it is inappropriate to indiscriminately apply PRP through every stage of a contemporary IR system. Such systems contain multiple stages (e.g., retrieval, pre-ranking, ranking, and re-ranking stages, as examined in this paper). The selection bias inherent in the model of each stage significantly influences the results that are ultimately presented to users. To address this issue, we propose an improved ranking principle for multi-stage systems, namely the Generalized Probability Ranking Principle (GPRP), to emphasize both the selection bias in each stage of the system pipeline as well as the underlying interest of users. We realize GPRP via a unified algorithmic framework named Full Stage Learning to Rank. Our core idea is to first estimate the selection bias in the subsequent stages and then learn a ranking model that best complies with the downstream modules' selection bias so as to deliver its top ranked results to the final ranked list in the system's output. We performed extensive experiment evaluations of our developed Full Stage Learning to Rank solution, using both simulations and online A/B tests in one of the leading short-video recommendation platforms. The algorithm is proved to be effective in both retrieval and ranking stages. Since deployed, the algorithm has brought consistent and significant performance gain to the platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64d107a6-2b17-4277-9bd8-b33f79489cd6Cited by top-tier papers8
- Killing Two Birds with One Stone: Unifying Retrieval and Ranking with a Single Generative Recommendation ModelLuankang Zhang, Kenan Song, Yi Quan Lee, Wei Guo et al.SIGIR 2025 · 5 citations
- Denoising Neural Reranker for Recommender SystemsWenyu Mao, Shuchang Liu, HailanYang, Xiaobei Wang et al.ICLR 2026 · 4 citations
- Comprehensive List Generation for Multi-Generator RerankingHailan Yang, Zhenyu Qi, Shuchang Liu, Xiaoyu Yang et al.SIGIR 2025 · 4 citations
- Online Two-Stage Submodular MaximizationIasonas Nikolaou, Miltiadis Stouras, Stratis Ioannidis, Evimaria TerziNeurIPS 2025 · 1 citation
- Learning Cascade Ranking as One NetworkYunli Wang, Zhen Zhang, Zhiqiang Wang, Zixuan Yang et al.ICML 2025
Builds on4
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 60 citations
- RankFlow: Joint Optimization of Multi-Stage Cascade Ranking Systems as FlowsJiarui Qin, Jiachen Zhu, Bo Chen, Zhirong Liu et al.SIGIR 2022 · 31 citations
- PairRank: Online Pairwise Learning to Rank by Divide-and-ConquerYiling Jia, Huazheng Wang, Stephen D. Guo, Hongning WangWWW 2021 · 24 citations
Related papers
- Multi-Level Interaction Reranking with User Behavior HistoryYunjia Xi, Weiwen Liu, Jieming Zhu, Xilong Zhao et al.SIGIR 2022 · 21 citations
- Both Supply and Precision: Sample Debias and Ranking Consistency Joint Learning for Large Scale Pre-Ranking SystemFeng Gao, Xin Zhou, Yinning Shao, Yue Wu et al.AAAI 2025 · 2 citations
- Measuring and Mitigating Item Under-Recommendation Bias in Personalized Ranking SystemsZiwei Zhu, Jianling Wang, James CaverleeSIGIR 2020 · 103 citations
- Whole Page Unbiased Learning to RankHaitao Mao, Lixin Zou, Yujia Zheng, Jiliang Tang et al.WWW 2024 · 6 citations
- MDP2 Forest: A Constrained Continuous Multi-dimensional Policy Optimization Approach for Short-video RecommendationSizhe Yu, Ziyi Liu, Shixiang Wan, Jia Zheng et al.KDD 2022 · 4 citations
