PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User Engagement
Wanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun, Shuchang Liu, Dong Zheng, Peng Jiang, Kun Gai, Bo An
摘要
Current advances in recommender systems have been remarkably successful in optimizing immediate engagement. However, long-term user engagement, a more desirable performance metric, remains difficult to improve. Meanwhile, recent reinforcement learning (RL) algorithms have shown their effectiveness in a variety of long-term goal optimization tasks. For this reason, RL is widely considered as a promising framework for optimizing long-term user engagement in recommendation. Though promising, the application of RL heavily relies on well-designed rewards, but designing rewards related to long-term user engagement is quite difficult. To mitigate the problem, we propose a novel paradigm, recommender systems with human preferences (or Preference-based Recommender systems), which allows RL recommender systems to learn from preferences about users' historical behaviors rather than explicitly defined rewards. Such preferences are easily accessible through techniques such as crowdsourcing, as they do not require any expert knowledge. With PrefRec, we can fully exploit the advantages of RL in optimizing long-term goals, while avoiding complex reward engineering. PrefRec uses the preferences to automatically train a reward function in an end-to-end manner. The reward function is then used to generate learning signals to train the recommendation policy. Furthermore, we design an effective optimization method for PrefRec, which uses an additional value function, expectile regression and reward model pre-training to improve the performance. We conduct experiments on a variety of long-term user engagement optimization tasks. The results show that PrefRec significantly outperforms previous state-of-the-art methods in all the tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecutionGengrui Zhang, Yao Wang, Xiaoshuang Chen, Hongyi Qian 等AAAI 2024 · 被引用 12 次
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 被引用 5 次
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li 等WWW 2025 · 被引用 4 次
- xMTF: A Formula-Free Model for Reinforcement-Learning-Based Multi-Task Fusion in Recommender SystemsYang Cao, Changhao Zhang, Xiaoshuang Chen, Kaiqiao Zhan 等WWW 2025 · 被引用 3 次
- DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic RewardYi Zhang, Ruihong Qiu, Xuwei Xu, Jiajun Liu 等SIGIR 2025 · 被引用 3 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Algorithmic Effects on the Diversity of Consumption on SpotifyAshton Anderson, Lucas Maystre, Ian Anderson, Rishabh Mehrotra 等WWW 2020 · 被引用 211 次
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee 等ICLR 2022 · 被引用 116 次
相关 Paper
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng 等ICLR 2023 · 被引用 6 次
- An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing PlatformsCaihua Shan, Nikos Mamoulis, Reynold Cheng, Guoliang Li 等ICDE 2020 · 被引用 23 次
- Towards Off-Policy Learning for Ranking Policies with Logged FeedbackTeng Xiao, Suhang WangAAAI 2022 · 被引用 8 次
- Implicit Safety Alignment from Crowd PreferencesQian Lin, Daniel S BrownICML 2026
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen 等WWW 2026
