Off-policy Learning in Two-stage Recommender Systems
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, Ed H. Chi
摘要
Many real-world recommender systems need to be highly scalable: matching millions of items with billions of users, with milliseconds latency. The scalability requirement has led to widely used two-stage recommender systems, consisting of efficient candidate generation model(s) in the first stage and a more powerful ranking model in the second stage. Logged user feedback, e.g., user clicks or dwell time, are often used to build both candidate generation and ranking models for recommender systems. While it’s easy to collect large amount of such data, they are inherently biased because the feedback can only be observed on items recommended by the previous systems. Recently, off-policy correction on such biases have attracted increasing interest in the field of recommender system research. However, most existing work either assumed that the recommender system is a single-stage system or only studied how to apply off-policy correction to the candidate generation stage of the system without explicitly considering the interactions between the two stages. In this work, we propose a two-stage off-policy policy gradient method, and showcase that ignoring the interaction between the two stages leads to a sub-optimal policy in two-stage recommender systems. The proposed method explicitly takes into account the ranking model when training the candidate generation model, which helps improve the performance of the whole system. We conduct experiments on real-world datasets with large item space and demonstrate the effectiveness of our proposed method.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper21
- A Model of Two Tales: Dual Transfer Learning Framework for Improved Long-tail Item RecommendationYin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi 等WWW 2021 · 被引用 124 次
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue 等WWW 2023 · 被引用 60 次
- On Component Interactions in Two-Stage Recommender SystemsJiri Hron, Karl Krauth, Michael I. Jordan, Niki KilbertusNeurIPS 2021 · 被引用 39 次
- Revisiting Injective Attacks on Recommender SystemsHaoyang Li, Shimin Di, Lei ChenNeurIPS 2022 · 被引用 26 次
- NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingYitu Wang, Shiyu Li, Qilin Zheng, Linghao Song 等ISCA 2024 · 被引用 26 次
相关 Paper
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky 等WWW 2020 · 被引用 123 次
- MGPolicy: Meta Graph Enhanced Off-policy Learning for RecommendationsXiangmeng Wang, Qian Li, Dianer Yu, Zhichao Wang 等SIGIR 2022 · 被引用 8 次
- Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage RankingHaruka Kiyohara, Mihaela Curmei, Ariel Evnine, Shankar Kalyanaraman 等ICML 2026
- Off-policy Learning over Heterogeneous Information for RecommendationXiangmeng Wang, Qian Li, Dianer Yu, Guandong XuWWW 2022 · 被引用 11 次
- GoalRank: Group-Relative Optimization for a Large Ranking ModelKaike Zhang, Xiaobei Wang, Shuchang Liu, HailanYang 等ICLR 2026 · 被引用 2 次
