Large Language Models are Learnable Planners for Long-Term Recommendation
Wentao Shi, Xiangnan He, Yang Zhang, Chongming Gao, Xinyue Li, Jizhi Zhang, Qifan Wang, Fuli Feng
摘要
Planning for both immediate and long-term benefits becomes increasingly important in recommendation. Existing methods apply Reinforcement Learning (RL) to learn planning capacity by maximizing cumulative reward for long-term recommendation. However, the scarcity of recommendation data presents challenges such as instability and susceptibility to overfitting when training RL models from scratch, resulting in sub-optimal performance. In this light, we propose to leverage the remarkable planning capabilities over sparse data of Large Language Models (LLMs) for long-term recommendation. The key to achieving the target lies in formulating a guidance plan following principles of enhancing long-term engagement and grounding the plan to effective and executable actions in a personalized manner. To this end, we propose a Bi-level Learnable LLM Planner framework, which consists of a set of LLM instances and breaks down the learning process into macro-learning and micro-learning to learn macro-level guidance and micro-level personalized recommendation policies, respectively. Extensive experiments validate that the framework facilitates the planning ability of LLMs for long-term recommendation. Our code and data can be found at https://github.com/jizhi-zhang/BiLLP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- NextQuill: Causal Preference Modeling for Enhancing LLM PersonalizationXiaoyan Zhao, Juntao You, Yang Zhang, Wenjie Wang 等ICLR 2026 · 被引用 38 次
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form GenerationChengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao 等ICLR 2026 · 被引用 35 次
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng 等SIGIR 2025 · 被引用 15 次
- Interactive Recommendation Agent with Active User CommandsJiakai Tang, Wen Chen, Yujie Luo, Xunke Xi 等KDD 2026 · 被引用 14 次
- Agentic Feedback Loop Modeling Improves Recommendation and User SimulationShihao Cai, Jizhi Zhang, Keqin Bao, Chongming Gao 等SIGIR 2025 · 被引用 13 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin 等AAAI 2024 · 被引用 484 次
相关 Paper
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen 等NeurIPS 2025 · 被引用 3 次
- LLM4RSR: Large Language Models as Data Correctors for Robust Sequential RecommendationYatong Sun, Xiaochun Yang, Zhu Sun, Yan Wang 等AAAI 2025 · 被引用 2 次
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao 等SIGIR 2020 · 被引用 122 次
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen 等WWW 2026
- Factorized Latent Reasoning for LLM-based RecommendationTianqi Gao, Chengkai Huang, Zihan Wang, Cao Liu 等SIGIR 2026
