Causal Decision Transformer for Recommender Systems via Offline Reinforcement Learning
Siyu Wang, Xiaocong Chen, Dietmar Jannach, Lina Yao
Abstract
Reinforcement learning-based recommender systems have recently gained popularity. However, the design of the reward function, on which the agent relies to optimize its recommendation policy, is often not straightforward. Exploring the causality underlying users' behavior can take the place of the reward function in guiding the agent to capture the dynamic interests of users. Moreover, due to the typical limitations of simulation environments (e.g., data ineffi- ciency), most of the work cannot be broadly applied in large-scale situations. Although some works attempt to convert the offline dataset into a simulator, data inefficiency makes the learning pro- cess even slower. Because of the nature of reinforcement learning (i.e., learning by interaction), it cannot collect enough data to train during a single interaction. Furthermore, traditional reinforcement learning algorithms do not have a solid capability like supervised learning methods to learn from offline datasets directly. In this paper, we propose a new model named the causal decision transformer for recommender systems (CDT4Rec). CDT4Rec is an offline reinforce- ment learning system that can learn from a dataset rather than from online interaction. Moreover, CDT4Rec employs the transformer architecture, which is capable of processing large offline datasets and capturing both short-term and long-term dependencies within the data to estimate the causal relationship between action, state, and reward. To demonstrate the feasibility and superiority of our model, we have conducted experiments on six real-world offline datasets and one online simulator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cd03774-1675-4362-aa85-b34e8d2f5699Cited by top-tier papers6
- Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic RecommendationXiaocong Chen, Siyu Wang, Lina YaoKDD 2024 · 6 citations
- Policy-Guided Causal State Representation for Offline Reinforcement Learning RecommendationSiyu Wang, Xiaocong Chen, Lina YaoWWW 2025 · 5 citations
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement LearningChen Gong, Zheng Liu, Kecen Li, Tianhao WangNDSS 2026 · 3 citations
- A General Framework for Off-Policy Learning with Partially-Observed RewardRikiya Takehi, Masahiro Asami, Kosuke Kawakami, Yuta SaitoICLR 2025
- Reward-Preserving Counterfactual State Editing for Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Mingming Gong, Yong Li et al.ICML 2026
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
Related papers
- Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term RetentionZiru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai et al.SIGIR 2024 · 23 citations
- When should we prefer Decision Transformers for Offline Reinforcement Learning?Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani et al.ICLR 2024 · 18 citations
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
- User Retention-oriented Recommendation with Decision TransformerKesen Zhao, Lixin Zou, Xiangyu Zhao, Maolin Wang et al.WWW 2023 · 38 citations
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto et al.ICLR 2023 · 10 citations
