Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
Siwei Wang, Yifei Shen, Haoran Sun, Shi Feng, Shang-Hua Teng, Li Dong, Yaru Hao, Wei Chen
摘要
Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness remains elusive. In this work, we investigate RL's benefits and limitations through a tractable graph-based abstraction, focusing on policy gradient (PG) and Q-learning methods. Our theoretical analyses reveal that supervised fine-tuning (SFT) may introduce co-occurrence-based spurious solutions, whereas RL achieves correct planning primarily through exploration, underscoring exploration’s role in enabling better generalization. However, we also show that PG suffers from diversity collapse, where output diversity decreases during training and persists even after perfect accuracy is attained. By contrast, Q-learning provides two key advantages: off-policy learning and diversity preservation at convergence. We further demonstrate that careful reward design is necessary to prevent Q-value bias in Q-learning. Finally, applying our framework to the real-world planning benchmark Blocksworld, we confirm that these behaviors manifest in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient Estimation of Kernel Surrogate Models for Task AttributionZhenshuo Zhang, Minxuan Duan, Hongyang R. ZhangICLR 2026 · 被引用 6 次
- When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path CompressionXinnan Dai, Kai Yang, cheng Luo, Shenglai Zeng 等ICML 2026
它引用的顶会 Paper13
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- LLaGA: Large Language and Graph AssistantRunjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah 等ICML 2024 · 被引用 180 次
- Evaluating Cognitive Maps and Planning in Large Language Models with CogEvalIda Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma 等NeurIPS 2023 · 被引用 114 次
- Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics TasksMurtaza Dalal, Tarun Chiruvolu, Devendra Singh Chaplot, Ruslan SalakhutdinovICLR 2024 · 被引用 86 次
- Understanding Transformer Reasoning Capabilities via Graph AlgorithmsClayton Sanford, Bahare Fatemi, Ethan Hall, Anton Tsitsulin 等NeurIPS 2024 · 被引用 84 次
相关 Paper
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei 等ICLR 2026 · 被引用 168 次
- Differential Smoothing Mitigates Sharpening and Improves LLM ReasoningJingchu Gai, Guanning Zeng, Huaqing ZHANG, Aditi RaghunathanICML 2026 · 被引用 13 次
- Imitating Language via Scalable Inverse Reinforcement LearningMarkus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja 等NeurIPS 2024 · 被引用 26 次
