Locality-Sensitive State-Guided Experience Replay Optimization for Sparse Rewards in Online Recommendation
Xiaocong Chen, Lina Yao, Julian J. McAuley, Weili Guan, Xiaojun Chang, Xianzhi Wang
Abstract
Online recommendation requires handling rapidly changing user preferences. Deep reinforcement learning (DRL) is an effective means of capturing users' dynamic interest during interactions with recommender systems. Generally, it is challenging to train a DRL agent in online recommender systems because of the sparse rewards caused by the large action space (e.g., candidate item space) and comparatively fewer user interactions. Leveraging experience replay (ER) has been extensively studied to conquer the issue of sparse rewards. However, they adapt poorly to the complex environment of online recommender systems and are inefficient in learning an optimal strategy from past experience. As a step to filling this gap, we propose a novel state-aware experience replay model, in which the agent selectively discovers the most relevant and salient experiences and is guided to find the optimal policy for online recommendations. In particular, a locality-sensitive hashing method is proposed to selectively retain the most meaningful experience at scale and a prioritized reward-driven strategy is designed to replay more valuable experiences with higher chance. We formally show that the proposed method guarantees the upper and lower bound on experience replay and optimizes the space complexity, as well as empirically demonstrate our model's superiority to several existing experience replay methods over three benchmark simulation platforms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic RecommendationXiaocong Chen, Siyu Wang, Lina YaoKDD 2024 · 6 citations
- Contrastive Representation for Interactive RecommendationJingyu Li, Zhiyong Feng, Dongxiao He, Hongqi Chen et al.AAAI 2025 · 2 citations
Builds on3
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Leveraging Demonstrations for Reinforcement Recommendation Reasoning over Knowledge GraphsKangzhi Zhao, Xiting Wang, Yuren Zhang, Li Zhao et al.SIGIR 2020 · 114 citations
- Attentive Experience ReplayPeiquan Sun, Wengang Zhou, Houqiang LiAAAI 2020 · 62 citations
Related papers
- Prioritized Generative ReplayRenhao Wang, Kevin Frans, Pieter Abbeel, Sergey Levine et al.ICLR 2025
- Reliability-Adjusted Prioritized Experience ReplayLeonard S. Pleiss, Tobias Sutter, Maximilian SchifferICLR 2026 · 3 citations
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang et al.AAAI 2021 · 131 citations
- The Benefits of Model-Based Generalization in Reinforcement LearningKenny John Young, Aditya A. Ramesh, Louis Kirsch, Jürgen SchmidhuberICML 2023 · 18 citations
- RLPer: A Reinforcement Learning Model for Personalized SearchJing Yao, Zhicheng Dou, Jun Xu, Ji-Rong WenWWW 2020 · 33 citations
