Rethinking Reinforcement Learning for Recommendation: A Prompt Perspective
Xin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren, Konstantina Christakopoulou, Zhaochun Ren
摘要
Modern recommender systems aim to improve user experience. As reinforcement learning (RL) naturally fits this objective-maximizing an user's reward per session-it has become an emerging topic in recommender systems. Developing RL-based recommendation methods, however, is not trivial due to the offline training challenge. Specifically, the keystone of traditional RL is to train an agent with large amounts of online exploration making lots of 'errors' in the process. In the recommendation setting, though, we cannot afford the price of making 'errors' online. As a result, the agent needs to be trained through offline historical implicit feedback, collected under different recommendation policies; traditional RL algorithms may lead to sub-optimal policies under these offline training settings.
Here we propose a new learning paradigm-namely Prompt-Based Reinforcement Learning (PRL)-for the offline training of RL-based recommendation agents. While traditional RL algorithms attempt to map state-action input pairs to their expected rewards (e.g., Q-values), PRL directly infers actions (i.e., recommended items) from state-reward inputs. In short, the agents are trained to predict a recommended item given the prior interactions and an observed reward value-with simple supervised learning. At deployment time, this historical (training) data acts as a knowledge base, while the state-reward pairs are used as a prompt. The agents are thus used to answer the question: Which item should be recommended given the prior interactions & the prompted reward value? We implement PRL with four notable recommendation models and conduct experiments on two real-world e-commerce datasets. Experimental results demonstrate the superior performance of our proposed methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Denoising and Prompt-Tuning for Multi-Behavior RecommendationChi Zhang, Rui Chen, Xiangyu Zhao, Qilong Han 等WWW 2023 · 被引用 71 次
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang 等SIGIR 2023 · 被引用 65 次
- PromptMM: Multi-Modal Knowledge Distillation for Recommendation with Prompt-TuningWei Wei, Jiabin Tang, Lianghao Xia, Yangqin Jiang 等WWW 2024 · 被引用 46 次
- User Retention-oriented Recommendation with Decision TransformerKesen Zhao, Lixin Zou, Xiangyu Zhao, Maolin Wang 等WWW 2023 · 被引用 38 次
- Reinforcement Learning-based Recommender Systems with Large Language Models for State Reward and Action ModelingJie Wang, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2024 · 被引用 27 次
它引用的顶会 Paper5
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 被引用 950 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 被引用 217 次
相关 Paper
- Contrastive State Augmentations for Reinforcement Learning-Based Recommender SystemsZhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren 等SIGIR 2023 · 被引用 20 次
- Preference Elicitation for Offline Reinforcement LearningAlizée Pace, Bernhard Schölkopf, Gunnar Rätsch, Giorgia RamponiICLR 2025
- Text-Based Interactive Recommendation via Offline Reinforcement LearningRuiyi Zhang, Tong Yu, Yilin Shen, Hongxia JinAAAI 2022 · 被引用 10 次
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 被引用 82 次
- Causal Decision Transformer for Recommender Systems via Offline Reinforcement LearningSiyu Wang, Xiaocong Chen, Dietmar Jannach, Lina YaoSIGIR 2023 · 被引用 33 次
