User Retention-oriented Recommendation with Decision Transformer
Kesen Zhao, Lixin Zou, Xiangyu Zhao, Maolin Wang, Dawei Yin
Abstract
Improving user retention with reinforcement learning (RL) has attracted increasing attention due to its significant importance in boosting user engagement. However, training the RL policy from scratch without hurting users' experience is unavoidable due to the requirement of trial-and-error searches. Furthermore, the offline methods, which aim to optimize the policy without online interactions, suffer from the notorious stability problem in value estimation or unbounded variance in counterfactual policy evaluation. To this end, we propose optimizing user retention with Decision Transformer (DT), which avoids the offline difficulty by translating the RL as an autoregressive problem. However, deploying the DT in recommendation is a non-trivial problem because of the following challenges: (1) deficiency in modeling the numerical reward value; (2) data discrepancy between the policy learning and recommendation generation; (3) unreliable offline performance evaluation. In this work, we, therefore, contribute a series of strategies for tackling the exposed issues. We first articulate an efficient reward prompt by weighted aggregation of meta embeddings for informative reward embedding. Then, we endow a weighted contrastive learning method to solve the discrepancy between training and inference. Furthermore, we design two robust offline metrics to measure user retention. Finally, the significant improvement in the benchmark datasets demonstrates the superiority of the proposed method. The implementation code is available at https://github.com/kesenzhao/DT4Rec.git . CCS CONCEPTS • Information Systems → Recommender Systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3324fccf-f19a-4997-8c72-dcdee658bf0dCited by top-tier papers9
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao et al.SIGIR 2023 · 86 citations
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu et al.WWW 2025 · 50 citations
- Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term RetentionZiru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai et al.SIGIR 2024 · 23 citations
- STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential RecommendationMaolin Wang, Sheng Zhang, Ruocheng Guo, Wanyu Wang et al.SIGIR 2025 · 12 citations
- LLM-Powered User Simulator for Recommender SystemZijian Zhang, Shuchang Liu, Ziru Liu, Rui Zhong et al.AAAI 2025 · 11 citations
Builds on5
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Neural Interactive Collaborative FilteringLixin Zou, Long Xia, Yulong Gu, Xiangyu Zhao et al.SIGIR 2020 · 121 citations
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
- A Reinforcement Learning Framework for Relevance FeedbackAli Montazeralghaem, Hamed Zamani, James AllanSIGIR 2020 · 38 citations
- UserSim: User Simulation via Supervised GenerativeAdversarial NetworkXiangyu Zhao, Long Xia, Lixin Zou, Hui Liu et al.WWW 2021 · 31 citations
Related papers
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 33 citations
- Maximum-Entropy Regularized Decision Transformer with Reward Relabelling for Dynamic RecommendationXiaocong Chen, Siyu Wang, Lina YaoKDD 2024 · 6 citations
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- AURO: Reinforcement Learning for Adaptive User Retention Optimization in Recommender SystemsZhenghai Xue, Qingpeng Cai, Bin Yang, Lantao Hu et al.WWW 2025 · 6 citations
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng et al.ICLR 2023 · 6 citations
