Off-policy Learning in Two-stage Recommender Systems
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, Ed H. Chi
Abstract
Many real-world recommender systems need to be highly scalable: matching millions of items with billions of users, with milliseconds latency. The scalability requirement has led to widely used two-stage recommender systems, consisting of efficient candidate generation model(s) in the first stage and a more powerful ranking model in the second stage. Logged user feedback, e.g., user clicks or dwell time, are often used to build both candidate generation and ranking models for recommender systems. While it’s easy to collect large amount of such data, they are inherently biased because the feedback can only be observed on items recommended by the previous systems. Recently, off-policy correction on such biases have attracted increasing interest in the field of recommender system research. However, most existing work either assumed that the recommender system is a single-stage system or only studied how to apply off-policy correction to the candidate generation stage of the system without explicitly considering the interactions between the two stages. In this work, we propose a two-stage off-policy policy gradient method, and showcase that ignoring the interaction between the two stages leads to a sub-optimal policy in two-stage recommender systems. The proposed method explicitly takes into account the ranking model when training the candidate generation model, which helps improve the performance of the whole system. We conduct experiments on real-world datasets with large item space and demonstrate the effectiveness of our proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b70d88d0-dfe2-47f4-ada4-8539f7ebb898Cited by top-tier papers21
- A Model of Two Tales: Dual Transfer Learning Framework for Improved Long-tail Item RecommendationYin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi et al.WWW 2021 · 124 citations
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue et al.WWW 2023 · 60 citations
- On Component Interactions in Two-Stage Recommender SystemsJiri Hron, Karl Krauth, Michael I. Jordan, Niki KilbertusNeurIPS 2021 · 39 citations
- Revisiting Injective Attacks on Recommender SystemsHaoyang Li, Shimin Di, Lei ChenNeurIPS 2022 · 26 citations
- NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingYitu Wang, Shiyu Li, Qilin Zheng, Linghao Song et al.ISCA 2024 · 26 citations
Related papers
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- MGPolicy: Meta Graph Enhanced Off-policy Learning for RecommendationsXiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang et al.SIGIR 2022 · 8 citations
- Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage RankingHaruka Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman et al.ICML 2026
- Off-policy Learning over Heterogeneous Information for RecommendationXiangmeng Wang, Qian Li, Dianer Yu, Guandong XuWWW 2022 · 11 citations
- GoalRank: Group-Relative Optimization for a Large Ranking ModelKaike Zhang, Xiaobei Wang, Shuchang Liu, HailanYang et al.ICLR 2026 · 2 citations
