Reinforcement Learning with a Disentangled Universal Value Function for Item Recommendation
Kai Wang, Zhene Zou, Qilin Deng, Jianrong Tao, Runze Wu, Changjie Fan, Liang Chen, Peng Cui
Abstract
In recent years, there are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems (RS). In this paper, we summarize three key practical challenges of large-scale RL-based recommender systems: massive state and action spaces, high-variance environment, and the unspecific reward setting in recommendation. All these problems remain largely unexplored in the existing literature and make the application of RL challenging. We develop a model-based reinforcement learning framework, called GoalRec. Inspired by the ideas of world model (model-based), value function estimation (model-free), and goal-based RL, a novel disentangled universal value function designed for item recommendation is proposed. It can generalize to various goals that the recommender may have, and disentangle the stochastic environmental dynamics and high-variance reward signals accordingly. As a part of the value function, free from the sparse and high-variance reward signals, a high-capacity reward-independent world model is trained to simulate complex environmental dynamics under a certain goal. Based on the predicted environmental dynamics, the disentangled universal value function is related to the user's future trajectory instead of a monolithic state and a scalar reward. We demonstrate the superiority of GoalRec over previous approaches in terms of the above three practical challenges in a series of simulations and a real application.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29a66e15-e457-4dd9-8661-0abf2a89d8ecCited by top-tier papers1
Ask how each one uses itRelated papers
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li et al.WWW 2025 · 4 citations
- MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for RecommendationsDongyang Zhao, Liang Zhang, Bo Zhang, Lizhou Zheng et al.SIGIR 2020 · 33 citations
- The Adaptive Q-Network for Recommendation Tasks with Dynamic Item SpaceJianxiang Zhu, Dandan Lai, Zhongcui Ma, Yaxin PengAAAI 2025
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement LearningDeyao Zhu, Li Erran Li, Mohamed ElhoseinyICLR 2023
