Reinforcement Learning-based Recommender Systems with Large Language Models for State Reward and Action Modeling
Jie Wang, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. Jose
摘要
Reinforcement Learning (RL)-based recommender systems have demonstrated promising performance in session-based and sequential recommendation tasks. Existing offline RL-based sequential recommendation methods face the challenge of obtaining effective user feedback from the environment. Developing a model for the user state and shaping an appropriate reward for recommendation remains a challenge. In this paper, we leverage language understanding capabilities and adapt large language models (LLMs) as an environment (LE) to enhance RL-based recommenders. The LE is learned from a subset of user-item interaction data, thus reducing the need for large training data, and can synthesize user feedback for offline data by: (i) acting as a state model that produces high-quality states that enrich the user representation, and (ii) functioning as a reward model to accurately capture nuanced user preferences on actions. Moreover, the LE allows us to generate positive actions that augment the limited offline training data. We propose a LE Augmentation (LEA) method to further improve recommendation performance by optimising jointly the supervised component and the RL policy, using the augmented actions and historical user signals. We use LEA, the state, and reward models in conjunction with state-of-the-art RL recommenders and report experimental results on two publicly available datasets 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng 等SIGIR 2025 · 被引用 15 次
- CORONA: A Coarse-to-Fine Framework for Graph-based Recommendation with Large Language ModelsJunze Chen, Xinjie Yang, Cheng Yang, Junfei Bao 等SIGIR 2025 · 被引用 5 次
- Hierarchical Tree Search-based User Lifelong Behavior Modeling on Large Language ModelYu Xia, Rui Zhong, Hao Gu, Wei Yang 等SIGIR 2025 · 被引用 5 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan 等NeurIPS 2023 · 被引用 474 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li 等KDD 2022 · 被引用 245 次
相关 Paper
- Enhancing Sequential Recommenders with Augmented Knowledge from Aligned Large Language ModelsYankun Ren, Zhongde Chen, Xinxing Yang, Longfei Li 等SIGIR 2024 · 被引用 28 次
- Rethinking Reinforcement Learning for Recommendation: A Prompt PerspectiveXin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren 等SIGIR 2022 · 被引用 48 次
- Contrastive State Augmentations for Reinforcement Learning-Based Recommender SystemsZhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren 等SIGIR 2023 · 被引用 20 次
- Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?Sein Kim, Hongseok Kang, Kibum Kim, Jiwan Kim 等KDD 2025 · 被引用 3 次
- DELRec: Distilling Sequential Pattern to Enhance LLMs-Based Sequential RecommendationHaoyi Zhang, Guohao Sun, Jinhu Lu, Guanfeng Liu 等ICDE 2025 · 被引用 1 次
