Reinforcement Recommendation with User Multi-aspect Preference
Xu Chen, Yali Du, Long Xia, Jun Wang
摘要
Formulating recommender system with reinforcement learning (RL) frameworks has attracted increasing attention from both academic and industry communities. While many promising results have been achieved, existing models mostly simulate the environment reward with a unified value, which may hinder the understanding of users' complex preferences and limit the model performance. In this paper, we consider how to model user multi-aspect preferences in the context of RL-based recommender system. More specifically, we base our model on the framework of deterministic policy gradient (DPG), which is effective in dealing with large action spaces. A major challenge for modeling user multi-aspect preferences lies in the fact that they may contradict with each other. To solve this problem, we introduce Pareto optimization into the DPG framework. We assign each aspect with a tailored critic, and all the critics share the same actor. The Pareto optimization is realized by a gradient-based method, which can be easily integrated into the actor and critic learning process. Based on the designed model, we theoretically analyze its gradient bias in the optimization process, and we design a weight-reuse mechanism to lower the upper bound of this bias, which is shown to be effective for improving the model performance. We conduct extensive experiments based on three real-world datasets to demonstrate our model's superiorities. CCS CONCEPTS • Information systems → Personalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue 等WWW 2023 · 被引用 60 次
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 被引用 21 次
- Shilling Black-box Review-based Recommender Systems through Fake Review GenerationHung-Yun Chiang, Yi-Syuan Chen, Yun-Zhu Song, Hong-Han Shuai 等KDD 2023 · 被引用 15 次
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu 等ICML 2024 · 被引用 4 次
- How to Find the Exact Pareto Front for Multi-Objective MDPs?Yining Li, Peizhong Ju, Ness B. ShroffICLR 2025
它引用的顶会 Paper1
相关 Paper
- Personalized Approximate Pareto-Efficient RecommendationRuobing Xie, Yanlei Liu, Shaoliang Zhang, Rui Wang 等WWW 2021 · 被引用 46 次
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- Distributional Pareto-Optimal Multi-Objective Reinforcement LearningXin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian 等NeurIPS 2023 · 被引用 46 次
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue 等ACL 2025 · 被引用 13 次
- Multi-Task Recommendations with Reinforcement LearningZiru Liu, Jiejie Tian, Qingpeng Cai, Xiangyu Zhao 等WWW 2023 · 被引用 57 次
