HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM Agents
Jiangweizhi Peng, Yuanxin Liu, Ruida Zhou, Charles Fleming, Zhaoran Wang, Alfredo Garcia, Mingyi Hong
Abstract
Training LLMs as interactive agents for multi-turn decision-making remains challenging, particularly in long-horizon tasks with sparse and delayed rewards, where agents must execute extended sequences of actions before receiving meaningful feedback. Most existing reinforcement learning (RL) methods model LLM agents as flat policies operating at a single time scale, selecting an action at each turn. In sparse-reward settings, this forces the agent to infer long-range dependencies solely from distant end-of-trajectory signals, often leading to inefficient learning and unstable behavior in complex environments. We propose HiPER , a novel Hi erarchical P lan– E xecute R L framework that jointly models and optimizes high-level subgoal planning and low-level action execution for LLM agents to overcome flat RL's brittle long-horizon behavior and weak credit assignment under sparse outcome feedback. By maintaining persistent subgoals across multiple turns and explicitly deciding when to switch between them, HiPER introduces structured intermediate decision points that facilitate learning under sparse feedback, converting implicit multi-turn structure into learnable decisions at different time scales. To enable effective training, we introduce Hierarchical Advantage Estimation (HAE), a two-timescale policy gradient method that assigns credit to both action execution and subgoal transitions and achieves variance reduction relative to flat advantage estimation. Empirically, HiPER achieves state-of-the-art performance on challenging interactive benchmarks, reaching 97.4% success on ALFWorld (+6.6% over the best prior method) and 83.3% on WebShop, with especially large gains on long-horizon tasks requiring multiple dependent subtasks. These results highlight the importance of explicit hierarchical decomposition for scalable RL training of multi-turn LLM agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c8401ff-83d1-461f-b60c-9e367ab55536Builds on4
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang et al.ICLR 2026 · 46 citations
- Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMsArash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee et al.ACL 2024 · 20 citations
Related papers
- Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningZican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue et al.ICML 2025
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan et al.ICML 2026 · 8 citations
- SHADOW: Dynamic-Aware Credit Assignment Against Long-Horizon TasksYuze Liu, Chaochao Lu, Chao YangAAAI 2026
- Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM AgentsShuai Zhen, Yanhua Yu, Ruopei Guo, Nan Cheng et al.ACL 2026 · 2 citations
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM AgentsJiawei Wang, Jiacai Liu, Yuqian Fu, Yingru Li et al.ICML 2026 · 38 citations
