AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender Systems
Zhenghai Xue, Qingpeng Cai, Bin Yang, Lantao Hu, Peng Jiang, Kun Gai, Bo An
摘要
The field of Reinforcement Learning (RL) has garnered increasing attention for its ability of optimizing user retention in recommender systems. A primary obstacle in this optimization process is the environment non-stationarity stemming from the continual and complex evolution of user behavior patterns over time, such as variations in interaction rates and retention propensities. These changes pose significant challenges to existing RL algorithms for recommendations, leading to issues with dynamics and reward distribution shifts. This paper introduces a novel approach called Adaptive User Retention Optimization (AURO) to address this challenge. To navigate the recommendation policy in non-stationary environments, AURO introduces an state abstraction module in the policy network. The module is trained with a new value-based loss function, aligning its output with the estimated performance of the current policy. As the policy performance of RL is sensitive to environment drifts, the loss function enables the state abstraction to be reflective of environment changes and notify the recommendation policy to adapt accordingly. Additionally, the non-stationarity of the environment introduces the problem of implicit cold start, where the recommendation policy continuously interacts with users displaying novel behavior patterns. AURO encourages exploration guarded by performance-based rejection sampling to maintain a stable recommendation quality in the cost-sensitive online environment. Extensive empirical analysis are conducted in a user retention simulator, the MovieLens dataset, and a live short-video recommendation platform, demonstrating AURO's superior performance against all evaluated baseline algorithms. Code is available at https://github.com/AIDefender/AURO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee 等ICML 2020 · 被引用 158 次
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 被引用 138 次
- Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) OptimismWang Chi Cheung, David Simchi-Levi, Ruihao ZhuICML 2020 · 被引用 114 次
相关 Paper
- User Retention-oriented Recommendation with Decision TransformerKesen Zhao, Lixin Zou, Xiangyu Zhao, Maolin Wang 等WWW 2023 · 被引用 38 次
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov 等ICLR 2025
- The Adaptive Q-Network for Recommendation Tasks with Dynamic Item SpaceJianxiang Zhu, Dandan Lai, Zhongcui Ma, Yaxin PengAAAI 2025
- Tackling Non-Stationarity in Reinforcement Learning via Causal-Origin RepresentationWanpeng Zhang, Yilin Li, Boyu Yang, Zongqing LuICML 2024 · 被引用 5 次
- Neural Interactive Collaborative FilteringLixin Zou, Long Xia, Yulong Gu, Xiangyu Zhao 等SIGIR 2020 · 被引用 121 次
