HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM Agents
Jiangweizhi Peng, Yuanxin Liu, Ruida Zhou, Charles Fleming, Zhaoran Wang, Alfredo Garcia, Mingyi Hong
摘要
Training LLMs as interactive agents for multi-turn decision-making remains challenging, particularly in long-horizon tasks with sparse and delayed rewards, where agents must execute extended sequences of actions before receiving meaningful feedback. Most existing reinforcement learning (RL) methods model LLM agents as flat policies operating at a single time scale, selecting an action at each turn. In sparse-reward settings, this forces the agent to infer long-range dependencies solely from distant end-of-trajectory signals, often leading to inefficient learning and unstable behavior in complex environments. We propose HiPER , a novel Hi erarchical P lan– E xecute R L framework that jointly models and optimizes high-level subgoal planning and low-level action execution for LLM agents to overcome flat RL's brittle long-horizon behavior and weak credit assignment under sparse outcome feedback. By maintaining persistent subgoals across multiple turns and explicitly deciding when to switch between them, HiPER introduces structured intermediate decision points that facilitate learning under sparse feedback, converting implicit multi-turn structure into learnable decisions at different time scales. To enable effective training, we introduce Hierarchical Advantage Estimation (HAE), a two-timescale policy gradient method that assigns credit to both action execution and subgoal transitions and achieves variance reduction relative to flat advantage estimation. Empirically, HiPER achieves state-of-the-art performance on challenging interactive benchmarks, reaching 97.4% success on ALFWorld (+6.6% over the best prior method) and 83.3% on WebShop, with especially large gains on long-horizon tasks requiring multiple dependent subtasks. These results highlight the importance of explicit hierarchical decomposition for scalable RL training of multi-turn LLM agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 被引用 1,539 次
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang 等ICLR 2026 · 被引用 46 次
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee 等ACL 2024 · 被引用 20 次
相关 Paper
- Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningZican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue 等ICML 2025
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan 等ICML 2026 · 被引用 8 次
- SHADOW: Dynamic-Aware Credit Assignment Against Long-Horizon TasksYuze Liu, Chaochao Lu, Chao YangAAAI 2026
- Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsShuai Zhen, Yanhua Yu, Ruopei Guo, Nan Cheng 等ACL 2026 · 被引用 2 次
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM AgentsJiawei Wang, Jiacai Liu, Yuqian Fu, Yingru Li 等ICML 2026 · 被引用 38 次
