Looking into User's Long-term Interests through the Lens of Conservative Evidential Learning
Dingrong Wang, Krishna Prasad Neupane, Ervine Zheng, Qi Yu
摘要
Reinforcement learning (RL) provides an effective means to capture users' evolving preferences, leading to improved recommendation performance over time. However, existing RL approaches primarily rely on standard exploration strategies, which are less effective for a large item space with sparse reward signals given the limited interactions for most users. Therefore, they may not be able to learn the optimal policy that effectively captures user's evolving preferences and achieves the maximum expected reward over the long term. In this paper, we propose a novel evidential conservative Q-learning framework (ECQL) that learns an effective and conservative recommendation policy by integrating evidence-based uncertainty and conservative learning. ECQL conducts evidence-aware explorations to discover items that are located beyond current observations but reflect users' long-term interests. It offers an uncertainty-aware conservative view on policy evaluation to discourage deviating too much from users' current interests. Two central components of ECQL include a uniquely designed sequential state encoder and a novel conservative evidential-actor-critic (CEAC) module. The former generates the current state of the environment by aggregating historical information and a sliding window that contains the current user interactions as well as newly recommended items from RL exploration that may represent short and long-term interests respectively. The latter performs an evidence-based rating prediction by maximizing the conservative evidential Q-value and leverages an uncertainty-aware ranking score to explore the item space for a more diverse and valuable recommendation. Experiments on multiple real-world dynamic datasets demonstrate the state-of-the-art performance of ECQL and its capability to capture users' long-term interests.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- Contrastive Learning for Sequential RecommendationXu Xie, Fei Sun, Zhaoyang Liu, Shiwen Wu 等ICDE 2022 · 被引用 674 次
- Disentangled Self-Supervision in Sequential RecommendersJianxin Ma, Chang Zhou, Hongxia Yang, Peng Cui 等KDD 2020 · 被引用 223 次
- Exploration and Regularization of the Latent Action Space in RecommendationShuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang 等WWW 2023 · 被引用 54 次
相关 Paper
- Evidential Stochastic Differential Equations for Time-Aware Sequential RecommendationKrishna Prasad Neupane, Ervine Zheng, Qi YuNeurIPS 2024 · 被引用 3 次
- The Adaptive Q-Network for Recommendation Tasks with Dynamic Item SpaceJianxiang Zhu, Dandan Lai, Zhongcui Ma, Yaxin PengAAAI 2025
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding 等SIGIR 2021 · 被引用 131 次
- Neural Interactive Collaborative FilteringLixin Zou, Long Xia, Yulong Gu, Xiangyu Zhao 等SIGIR 2020 · 被引用 121 次
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao 等SIGIR 2020 · 被引用 122 次
