ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor
Wanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng, Peng Jiang, Kun Gai, Bo An
摘要
Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dwell time. Meanwhile, reinforcement learning (RL) is widely regarded as a promising framework for optimizing long-term engagement in sequential recommendation. However, due to expensive online interactions, it is very difficult for RL algorithms to perform state-action value estimation, exploration and feature extraction when optimizing long-term engagement. In this paper, we propose ResAct which seeks a policy that is close to, but better than, the online-serving policy. In this way, we can collect sufficient data near the learned policy so that state-action values can be properly estimated, and there is no need to perform online interaction. ResAct optimizes the policy by first reconstructing the online behaviors and then improving it via a Residual Actor. To extract long-term information, ResAct utilizes two information-theoretical regularizers to confirm the expressiveness and conciseness of features. We conduct experiments on a benchmark dataset and a large-scale industrial dataset which consists of tens of millions of recommendation requests. Experimental results show that our method significantly outperforms the state-of-the-art baselines in various long-term engagement optimization tasks. * The work was done during an internship at Kuaishou Technology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang 等SIGIR 2023 · 被引用 65 次
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun 等KDD 2023 · 被引用 24 次
- UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecutionGengrui Zhang, Yao Wang, Xiaoshuang Chen, Hongyi Qian 等AAAI 2024 · 被引用 12 次
- Meta Clustering of Neural BanditsYikun Ban, Yunzhe Qi, Tianxin Wei, Lihui Liu 等KDD 2024 · 被引用 6 次
- AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender SystemsZhenghai Xue, Qingpeng Cai, Bin Yang, Lantao Hu 等WWW 2025 · 被引用 6 次
它引用的顶会 Paper3
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue 等WWW 2023 · 被引用 60 次
- MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for RecommendationsDongyang Zhao, Liang Zhang, Bo Zhang, Lizhou Zheng 等SIGIR 2020 · 被引用 33 次
相关 Paper
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao 等SIGIR 2020 · 被引用 122 次
- Lifelong Sequential Recommendation with Adaptive Subsequence Compression and Contextual FusionFei Li, Xiaoming Liu, Jiayi Luo, Guibing Guo 等WWW 2026
- Dynamic Memory based Attention Network for Sequential RecommendationQiaoyu Tan, Jianwei Zhang, Ninghao Liu, Xiao Huang 等AAAI 2021 · 被引用 74 次
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 被引用 217 次
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang 等WWW 2022 · 被引用 14 次
