OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
Zihao Wang, Muyao Li, Kaichen He, Xiangyu Wang, Zhancun Mu, Minghao Liu, Anji Liu, Yitao Liang
摘要
The choice of action representation is a critical yet unresolved challenge in developing capable, end-toend trainable agents. This paper first presents a large-scale, systematic comparison of prominent action abstractions for Vision-Language-Action (VLA) or hierarchical agent models in the open-ended Minecraft. Our analysis reveals that no single action space is universally optimal; instead, the most effective abstraction is highly task-dependent, creating a dilemma for building generalist agents. To resolve this, we introduce Chain of Action (CoA), a novel framework that unifies high-level planning and low-level control within a single, monolithic VLA model. CoA treats an abstracted action not as a command for a separate policy, but as an intermediate reasoning step-akin to a chain of thought-that guides the generation of the final, executable action. Furthermore, we demonstrate that an All-in-One agent trained on a diverse mixture of action spaces using the CoA paradigm learns a more robust and generalizable policy. This unified agent achieves a new state-of-the-art, improving the overall task success rate over strong, specialized baselines. To foster reproducible research, we release the OpenHA (Open Hierarchical Agents) suite, which includes our comprehensive benchmark of over 800 distinct tasks, curated datasets, source code, and all pretrained model checkpoints at: https://github.com/CraftJarvis/OpenHA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory SketchesJiayuan Gu, Sean Kirmani, Paul Wohlhart, Yao Lu 等ICLR 2024 · 被引用 135 次
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba 等NeurIPS 2023 · 被引用 123 次
相关 Paper
- OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following AgentsZihao Wang, Shaofei Cai, Zhancun Mu, Haowei Lin 等NeurIPS 2024 · 被引用 37 次
- Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationWenbo Zhang, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao 等NeurIPS 2025 · 被引用 19 次
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsQingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu 等CVPR 2025
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous InstructionsHelong Huang, Min Cen, Kai Tan, Xingyue Quan 等AAAI 2026 · 被引用 12 次
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong 等CVPR 2026 · 被引用 43 次
