ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor
Wanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng, Peng Jiang, Kun Gai, Bo An
Abstract
Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dwell time. Meanwhile, reinforcement learning (RL) is widely regarded as a promising framework for optimizing long-term engagement in sequential recommendation. However, due to expensive online interactions, it is very difficult for RL algorithms to perform state-action value estimation, exploration and feature extraction when optimizing long-term engagement. In this paper, we propose ResAct which seeks a policy that is close to, but better than, the online-serving policy. In this way, we can collect sufficient data near the learned policy so that state-action values can be properly estimated, and there is no need to perform online interaction. ResAct optimizes the policy by first reconstructing the online behaviors and then improving it via a Residual Actor. To extract long-term information, ResAct utilizes two information-theoretical regularizers to confirm the expressiveness and conciseness of features. We conduct experiments on a benchmark dataset and a large-scale industrial dataset which consists of tens of millions of recommendation requests. Experimental results show that our method significantly outperforms the state-of-the-art baselines in various long-term engagement optimization tasks. * The work was done during an internship at Kuaishou Technology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang et al.SIGIR 2023 · 65 citations
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun et al.KDD 2023 · 24 citations
- UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecutionGengrui Zhang, Yao Wang, Xiaoshuang Chen, Hongyi Qian et al.AAAI 2024 · 12 citations
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu et al.KDD 2024 · 6 citations
- AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender SystemsZhenghai Xue, Qingpeng Cai, Bin Yang, Lantao Hu et al.WWW 2025 · 6 citations
Builds on3
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue et al.WWW 2023 · 60 citations
- MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for RecommendationsDongyang Zhao, Liang Zhang, Bo Zhang, Lizhou Zheng et al.SIGIR 2020 · 33 citations
Related papers
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao et al.SIGIR 2020 · 122 citations
- Lifelong Sequential Recommendation with Adaptive Subsequence Compression and Contextual FusionFei Li, Xiaoming Liu, Jiayi Luo, Guibing Guo et al.WWW 2026
- Dynamic Memory based Attention Network for Sequential RecommendationQiaoyu Tan, Jianwei Zhang, Ninghao Liu, Xiao Huang et al.AAAI 2021 · 74 citations
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang et al.WWW 2022 · 14 citations
