Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic Recommendation
Xiaocong Chen, Siyu Wang, Lina Yao
Abstract
Reinforcement learning-based recommender systems have recently gained popularity. However, due to the typical limitations of simulation environments (e.g., data inefficiency), most of the work cannot be broadly applied in all domains. To counter these challenges, recent advancements have leveraged offline reinforcement learning methods, notable for their data-driven approach utilizing offline datasets. A prominent example of this is the Decision Transformer. Despite its popularity, the Decision Transformer approach has inherent drawbacks, particularly evident in recommendation methods based on it. This paper identifies two key shortcomings in existing Decision Transformer-based methods: a lack of stitching capability and limited effectiveness in online adoption. In response, we introduce a novel methodology named Max-Entropy enhanced Decision Transformer with Reward Relabeling for Offline RLRS (EDT4Rec). Our approach begins with a max entropy perspective, leading to the development of a max-entropy enhanced exploration strategy. This strategy is designed to facilitate more effective exploration in online environments. Additionally, to augment the model's capability to stitch sub-optimal trajectories, we incorporate a unique reward relabeling technique. To validate the effectiveness and superiority of EDT4Rec, we have conducted comprehensive experiments across six real-world offline datasets and in an online simulator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e1f5d76-9407-4477-a029-6ab3cdf7473aCited by top-tier papers2
- Policy-Guided Causal State Representation for Offline Reinforcement Learning RecommendationSiyu Wang, Xiaocong Chen, Lina YaoWWW 2025 · 5 citations
- Reward-Preserving Counterfactual State Editing for Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Mingming Gong, Yong Li et al.ICML 2026
Builds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Online Decision TransformerQinqing Zheng, Amy Zhang, Aditya GroverICML 2022 · 256 citations
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang et al.SIGIR 2023 · 65 citations
Related papers
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 33 citations
- Rethinking Decision Transformer via Hierarchical Reinforcement LearningYi Ma, Jianye Hao, Hebin Liang, Chenjun XiaoICML 2024 · 15 citations
- Reinformer: Max-Return Sequence Modeling for Offline RLZifeng Zhuang, Dengyun Peng, Jinxin Liu, Ziqi Zhang et al.ICML 2024 · 29 citations
- Less is More: an Attention-free Sequence Prediction Modeling for Offline Embodied LearningWei Huang, Jianshu Zhang, Leiyu Wang, Heyue Li et al.NeurIPS 2025
- Offline Reinforcement Learning with Adaptive Feature FusionTieru Wang, Kunbao Wu, Guoshun NanICLR 2026
