Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents
Shuai Zhen, Yanhua Yu, Ruopei Guo, Nan Cheng, Yang Deng
摘要
Large language model (LLM) agents have demonstrated strong capabilities in complex interactive decision-making tasks. However, existing LLM agents typically rely on increasingly long interaction histories, resulting in high computational cost and limited scalability. In this paper, we propose STEP-HRL, a hierarchical reinforcement learning (HRL) framework that enables step-level learning by conditioning only on single-step transitions rather than full interaction histories. STEP-HRL structures tasks hierarchically, using completed subtasks to represent global progress of overall task. By introducing a local progress module, it also iteratively and selectively summarizes interaction history within each subtask to produce a compact summary of local progress. Together, these components yield augmented step-level transitions for both high-level and low-level policies. Experimental results on ScienceWorld and ALFWorld benchmarks consistently demonstrate that STEP-HRL substantially outperforms baselines in terms of performance and generalization while reducing token usage. Our code is available at https: //github.com/TonyStark042/STEP-HRL . D Case Study Trajectory Comparison Normal Task Description: Your task is to find a(n) non-living thing. First, focus on the thing. Then, move it to the purple box in the bathroom. Observation: This room is called the kitchen. In it, you see the agent, a substance called air, a chair, a counter, a cupboard, a freezer, a fridge, a glass jar, a lighter, an oven, a painting, a sink, soap, a stopwatch, a stove, a table with a glass cup, and a thermometer. You also see: a door to the bathroom (closed), a door to the hallway (closed), and a door to the outside (closed). Action: open door to bathroom Reward: 0.08 Observation: The door is now open. Action: go to bathroom Reward: 0.17 Observation: This room is called the bathroom. In it, you see a bathtub, a drawing, a glass cup, a purple box, a sink, and a toilet. Action: focus on cup Reward: 0.50 Observation: You focus on the glass cup. Action: move cup to purple box Reward: 0.25 Observation: You move the glass cup to the purple box. done.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
相关 Paper
- HiPER: Hierarchical Plan–Execute RL for Multi-Turn LLM AgentsJiangweizhi Peng, Yuanxin Liu, Ruida Zhou, Charles Fleming 等ICML 2026
- Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningZican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue 等ICML 2025
- LightWM: Training-Free Hierarchical Working Memory for Small Language Model AgentsZiyi Wang, Haonan Jin, Zian Wang, Wendong Wang 等ICML 2026
- EPO: Hierarchical LLM Agents with Environment Preference OptimizationQi Zhao, Haotian Fu, Chen Sun, George KonidarisEMNLP 2024 · 被引用 3 次
- DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based AgentsJunShuo Zhang, Chengrui Huang, Feng Guo, Zihan Li 等ACL 2026
