Practical Counterfactual Policy Learning for Top-K Recommendations
Yaxu Liu, Jui-Nan Yen, Bo-Wen Yuan, Rundong Shi, Peng Yan, Chih-Jen Lin
摘要
For building recommender systems, a critical task is to learn a policy with collected feedback (e.g., ratings, clicks) to decide which items to be recommended to users. However, it has been shown that the selection bias in the collected feedback leads to biased learning and thus a sub-optimal policy. To deal with this issue, counterfactual learning has received much attention, where existing approaches can be categorized as either value learning or policy learning approaches. This work studies policy learning approaches for top-K recommendations with a large item space and points out several difficulties related to importance weight explosion, observation insufficiency, and training efficiency. A practical framework for policy learning is then proposed to overcome these difficulties. Our experiments confirm the effectiveness and efficiency of the proposed framework.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Policy-Aware Unbiased Learning to Rank for Top-k RankingsHarrie Oosterhuis, Maarten de RijkeSIGIR 2020 · 被引用 60 次
- Joint Policy-Value Learning for RecommendationOlivier Jeunen, David Rohde, Flavian Vasile, Martin BompaireKDD 2020 · 被引用 24 次
- Trustworthy Policy Learning under the Counterfactual No-Harm CriterionHaoxuan Li, Chunyuan Zheng, Yixiao Cao, Zhi Geng 等ICML 2023 · 被引用 34 次
- Treatment Effect Estimation for User Interest Exploration on Recommender SystemsJiaju Chen, Wenjie Wang, Chongming Gao, Peng Wu 等SIGIR 2024 · 被引用 8 次
- Off-policy Learning in Two-stage Recommender SystemsJiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang 等WWW 2020 · 被引用 106 次
