: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments
Sangeun Park, Minhae Kwon
摘要
A central goal of large language model (LLM) research is to build agentic systems that can plan, act, and adapt through sustained interaction with dynamic environments. While recent LLM-based agents exhibit impressive contextual reasoning, their long-horizon decision-making remains fragile, often suffering from , where goals and plans drift over extended interactions. We introduce , a hierarchical multi-agent decision-making framework that explicitly decomposes agent behavior into complementary roles. A high-level agent () focuses on context-aware sub-goal generation using supervised fine-tuning (SFT), while a low-level agent () executes atomic actions through offline-to-online reinforcement learning (RL) in interactive environments. This separation enables stable long-horizon control, mitigates objective drift, and allows efficient adaptation. Across diverse interactive environments, consistently outperforms strong agentic baselines, demonstrating improved robustness and coordination in multi-turn interaction. Beyond performance, we introduce and release three hierarchical benchmark datasets, filling a long-standing gap in training and evaluating hierarchical decision-making for LLM-based agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper43
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 被引用 326 次
相关 Paper
- CoPE: A Framework for Optimizing Coordination between Planning and Execution in LLM-based AgentsHuanxi Liu, Kun Hu, Qiang Wang, Yuanzhao Zhai 等ICML 2026
- ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model AgentsZhenyu Zhang, Tianyi Chen, Weiran Xu, Alex Pentland 等NeurIPS 2025 · 被引用 12 次
- DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn OptimizationJian Mu, Tianyi Lin, Chengwei Qin, Zhongxiang Dai 等ICML 2026
- Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement LearningZican Hu, Wei Liu, Xiaoye Qu, Xiangyu Yue 等ICML 2025
- Small LLMs Are Weak Tool Learners: A Multi-LLM AgentWeizhou Shen, Chenliang Li, Hongzhan Chen, Ming Yan 等EMNLP 2024 · 被引用 18 次
