On the Unexpected Effectiveness of Reinforcement Learning for Sequential Recommendation
Alvaro Labarca, Denis Parra, Rodrigo Toro Icarte
Abstract
In recent years, Reinforcement Learning (RL) has shown great promise in session-based recommendation. Sequential models that use RL have reached state-of-the-art performance for the Next-item Prediction (NIP) task. This result is intriguing, as the NIP task only evaluates how well the system can correctly recommend the next item to the user, while the goal of RL is to find a policy that optimizes rewards in the long term -sometimes at the expense of suboptimal shortterm performance. Then, how can RL improve the system's performance on short-term metrics? This article investigates this question by exploring proxy learning objectives, which we identify as goals RL models might be following, and thus could explain the performance boost. We found that RL -when used as an auxiliary loss -promotes the learning of embeddings that capture information about the user's previously interacted items. Subsequently, we replaced the RL objective with a straightforward auxiliary loss designed to predict the number of items the user interacted with. This substitution results in performance gains comparable to RL. These findings pave the way to improve performance and understanding of RL methods for recommender systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a753439-965a-4a43-9a3b-1fae19346e82Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- Leveraging Demonstrations for Reinforcement Recommendation Reasoning over Knowledge GraphsKangzhi Zhao, Xiting Wang, Yuren Zhang, Li Zhao et al.SIGIR 2020 · 114 citations
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
- Reinforced Anchor Knowledge Graph Generation for News Recommendation ReasoningDanyang Liu, Jianxun Lian, Zheng Liu, Xiting Wang et al.KDD 2021 · 49 citations
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren et al.SIGIR 2022 · 48 citations
Related papers
- Unsupervised Proxy Selection for Session-based Recommender SystemsJunsu Cho, SeongKu Kang, Dongmin Hyun, Hwanjo YuSIGIR 2021 · 21 citations
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao et al.SIGIR 2020 · 122 citations
- Incorporating User Micro-behaviors and Item Knowledge into Multi-task Learning for Session-based RecommendationWenjing Meng, Deqing Yang, Yanghua XiaoSIGIR 2020 · 122 citations
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng et al.ICLR 2023 · 6 citations
- NP-MiSR: Neural Process-based Multi-Interest Learning for Session-Based RecommendationJun Bao, Junbo Wang, Yiheng Jiang, Xiangfeng Liu et al.AAAI 2026
