Large Language Models are Learnable Planners for Long-Term Recommendation
Wentao Shi, Xiangnan He, Yang Zhang, Chongming Gao, Xinyue Li, Jizhi Zhang, Qifan Wang, Fuli Feng
Abstract
Planning for both immediate and long-term benefits becomes increasingly important in recommendation. Existing methods apply Reinforcement Learning (RL) to learn planning capacity by maximizing cumulative reward for long-term recommendation. However, the scarcity of recommendation data presents challenges such as instability and susceptibility to overfitting when training RL models from scratch, resulting in sub-optimal performance. In this light, we propose to leverage the remarkable planning capabilities over sparse data of Large Language Models (LLMs) for long-term recommendation. The key to achieving the target lies in formulating a guidance plan following principles of enhancing long-term engagement and grounding the plan to effective and executable actions in a personalized manner. To this end, we propose a Bi-level Learnable LLM Planner framework, which consists of a set of LLM instances and breaks down the learning process into macro-learning and micro-learning to learn macro-level guidance and micro-level personalized recommendation policies, respectively. Extensive experiments validate that the framework facilitates the planning ability of LLMs for long-term recommendation. Our code and data can be found at https://github.com/jizhi-zhang/BiLLP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73276e46-d022-41de-8237-656c8070a2a0Cited by top-tier papers17
- NextQuill: Causal Preference Modeling for Enhancing LLM PersonalizationXiaoyan Zhao, Juntao You, Yang Zhang, Wenjie Wang et al.ICLR 2026 · 38 citations
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form GenerationChengbing Wang, Yang Zhang, Wenjie Wang, Xiaoyan Zhao et al.ICLR 2026 · 35 citations
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng et al.SIGIR 2025 · 15 citations
- Interactive Recommendation Agent with Active User CommandsJiakai Tang, Wen Chen, Yujie Luo, Xunke Xi et al.KDD 2026 · 14 citations
- Agentic Feedback Loop Modeling Improves Recommendation and User SimulationShihao Cai, Jizhi Zhang, Keqin Bao, Chongming Gao et al.SIGIR 2025 · 13 citations
Builds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin et al.AAAI 2024 · 484 citations
Related papers
- Progress Reward Model for Reinforcement Learning via Large Language ModelsXiuhui Zhang, Ning Gao, Xingyu Jiang, Yihui Chen et al.NeurIPS 2025 · 3 citations
- LLM4RSR: Large Language Models as Data Correctors for Robust Sequential RecommendationYatong Sun, Xiaochun Yang, Zhu Sun, Yan Wang et al.AAAI 2025 · 2 citations
- KERL: A Knowledge-Guided Reinforcement Learning Model for Sequential RecommendationPengfei Wang, Yu Fan, Long Xia, Wayne Xin Zhao et al.SIGIR 2020 · 122 citations
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen et al.WWW 2026
- Factorized Latent Reasoning for LLM-based RecommendationTianqi Gao, Chengkai Huang, Zihan Wang, Cao Liu et al.SIGIR 2026
