Causal Decision Transformer for Recommender Systems via Offline Reinforcement Learning
Siyu Wang, Xiaocong Chen, Dietmar Jannach, Lina Yao
摘要
Reinforcement learning-based recommender systems have recently gained popularity. However, the design of the reward function, on which the agent relies to optimize its recommendation policy, is often not straightforward. Exploring the causality underlying users' behavior can take the place of the reward function in guiding the agent to capture the dynamic interests of users. Moreover, due to the typical limitations of simulation environments (e.g., data ineffi- ciency), most of the work cannot be broadly applied in large-scale situations. Although some works attempt to convert the offline dataset into a simulator, data inefficiency makes the learning pro- cess even slower. Because of the nature of reinforcement learning (i.e., learning by interaction), it cannot collect enough data to train during a single interaction. Furthermore, traditional reinforcement learning algorithms do not have a solid capability like supervised learning methods to learn from offline datasets directly. In this paper, we propose a new model named the causal decision transformer for recommender systems (CDT4Rec). CDT4Rec is an offline reinforce- ment learning system that can learn from a dataset rather than from online interaction. Moreover, CDT4Rec employs the transformer architecture, which is capable of processing large offline datasets and capturing both short-term and long-term dependencies within the data to estimate the causal relationship between action, state, and reward. To demonstrate the feasibility and superiority of our model, we have conducted experiments on six real-world offline datasets and one online simulator.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic RecommendationXiaocong Chen, Siyu Wang, Lina YaoKDD 2024 · 被引用 6 次
- Policy-Guided Causal State Representation for Offline Reinforcement Learning RecommendationSiyu Wang, Xiaocong Chen, Lina YaoWWW 2025 · 被引用 5 次
- PrivORL: Differentially Private Synthetic Dataset for Offline Reinforcement LearningChen Gong, Zheng Liu, Kecen Li, Tianhao WangNDSS 2026 · 被引用 3 次
- A General Framework for Off-Policy Learning with Partially-Observed RewardRikiya Takehi, Masahiro Asami, Kosuke Kawakami, Yuta SaitoICLR 2025
- Reward-Preserving Counterfactual State Editing for Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Mingming Gong, Yong Li 等ICML 2026
它引用的顶会 Paper7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 被引用 522 次
相关 Paper
- Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term RetentionZiru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai 等SIGIR 2024 · 被引用 23 次
- When should we prefer Decision Transformers for Offline Reinforcement Learning?Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani 等ICLR 2024 · 被引用 18 次
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren 等SIGIR 2022 · 被引用 48 次
- User Retention-oriented Recommendation with Decision TransformerKesen Zhao, Lixin Zou, Xiangyu Zhao, Maolin Wang 等WWW 2023 · 被引用 38 次
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto 等ICLR 2023 · 被引用 10 次
