PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User Engagement
Wanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun, Shuchang Liu, Dong Zheng, Peng Jiang, Kun Gai, Bo An
Abstract
Current advances in recommender systems have been remarkably successful in optimizing immediate engagement. However, long-term user engagement, a more desirable performance metric, remains difficult to improve. Meanwhile, recent reinforcement learning (RL) algorithms have shown their effectiveness in a variety of long-term goal optimization tasks. For this reason, RL is widely considered as a promising framework for optimizing long-term user engagement in recommendation. Though promising, the application of RL heavily relies on well-designed rewards, but designing rewards related to long-term user engagement is quite difficult. To mitigate the problem, we propose a novel paradigm, recommender systems with human preferences (or Preference-based Recommender systems), which allows RL recommender systems to learn from preferences about users' historical behaviors rather than explicitly defined rewards. Such preferences are easily accessible through techniques such as crowdsourcing, as they do not require any expert knowledge. With PrefRec, we can fully exploit the advantages of RL in optimizing long-term goals, while avoiding complex reward engineering. PrefRec uses the preferences to automatically train a reward function in an end-to-end manner. The reward function is then used to generate learning signals to train the recommendation policy. Furthermore, we design an effective optimization method for PrefRec, which uses an additional value function, expectile regression and reward model pre-training to improve the performance. We conduct experiments on a variety of long-term user engagement optimization tasks. The results show that PrefRec significantly outperforms previous state-of-the-art methods in all the tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84032e42-383a-43c2-be14-e9dab6dc4b24Cited by top-tier papers8
- UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecutionGengrui Zhang, Yao Wang, Xiaoshuang Chen, Hongyi Qian et al.AAAI 2024 · 12 citations
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic FrameworkKihyun Kim, Jiawei Zhang, Asuman Ozdaglar, Pablo A. ParriloICLR 2026 · 5 citations
- Value Function Decomposition in Markov Recommendation ProcessXiaobei Wang, Shuchang Liu, Qingpeng Cai, Xiang Li et al.WWW 2025 · 4 citations
- xMTF: A Formula-Free Model for Reinforcement-Learning-Based Multi-Task Fusion in Recommender SystemsYang Cao, Changhao Zhang, Xiaoshuang Chen, Kaiqiao Zhan et al.WWW 2025 · 3 citations
- DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic RewardYi Zhang, Ruihong Qiu, Xuwei Xu, Jiajun Liu et al.SIGIR 2025 · 3 citations
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Algorithmic Effects on the Diversity of Consumption on SpotifyAshton Anderson, Lucas Maystre, Ian Anderson, Rishabh Mehrotra et al.WWW 2020 · 211 citations
- SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement LearningJongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee et al.ICLR 2022 · 116 citations
Related papers
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng et al.ICLR 2023 · 6 citations
- An End-to-End Deep RL Framework for Task Arrangement in Crowdsourcing PlatformsCaihua Shan, Nikos Mamoulis, Reynold Cheng, Guoliang Li et al.ICDE 2020 · 23 citations
- Towards Off-Policy Learning for Ranking Policies with Logged FeedbackTeng Xiao, Suhang WangAAAI 2022 · 8 citations
- Implicit Safety Alignment from Crowd PreferencesQian Lin, Daniel S BrownICML 2026
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen et al.WWW 2026
