Value Function Decomposition in Markov Recommendation Process
Xiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li, Lantao Hu, Han Li, Guangming Xie
Abstract
Recent advances in recommender systems have shown that user-system interaction essentially formulates long-term optimization problems, and online reinforcement learning can be adopted to improve recommendation performance. The general solution framework incorporates a value function that estimates the user's expected cumulative rewards in the future and guides the training of the recommendation policy. To avoid local maxima, the policy may explore potential high-quality actions during inference to increase the chance of finding better future rewards. To accommodate the stepwise recommendation process, one widely adopted approach to learning the value function is learning from the difference between the values of two consecutive states of a user. However, we argue that this paradigm involves a challenge of Mixing Random Factors: there exist two random factors from the stochastic policy and the uncertain user environment, but they are not separately modeled in the standard temporal difference (TD) learning, which may result in a suboptimal estimation of the long-term rewards and less effective action exploration. As a solution, we show that these two factors can be separately approximated by decomposing the original temporal difference loss. The disentangled learning framework can achieve a more accurate estimation with faster learning and improved robustness against action exploration. As an empirical verification of our proposed method, we conduct offline experiments with simulated online environments built on the basis of public datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bcba1da6-7c2d-4fc3-9fd6-432815a974d7Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 138 citations
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang et al.AAAI 2021 · 131 citations
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue et al.WWW 2023 · 60 citations
- Multi-Task Recommendations with Reinforcement LearningZiru Liu, Jiejie Tian, Qingpeng Cai, Xiangyu Zhao et al.WWW 2023 · 57 citations
Related papers
- Reinforcement Learning with a Disentangled Universal Value Function for Item RecommendationKai Wang, Zhene Zou, Qilin Deng, Jianrong Tao et al.AAAI 2021 · 25 citations
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
- Prediction and Control in Continual Reinforcement LearningNishanth Anand, Doina PrecupNeurIPS 2023 · 26 citations
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska et al.ICML 2022 · 40 citations
- Advantage-Conditioned Flow Policy for Offline Reinforcement Learning in RecommendationXiaocong Chen, Siyu Wang, Lina YaoSIGIR 2026
