MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
Yiran Qin, Enshen Zhou, Qichang Liu, Zhenfei Yin, Lu Sheng, Ruimao Zhang, Yu Qiao, Jing Shao
Abstract
It is a long-lasting goal to design an embodied system that can solve long-horizon open-world tasks in human-like ways. However, existing approaches usually struggle with compound difficulties caused by the logic-aware decompo-sition and context-aware execution of these tasks. To this end, we introduce MP5, an open-ended multimodal em-bodied system built upon the challenging Minecraft sim-ulator, which can decompose feasible sub-objectives, de-sign sophisticated situation-aware plans, and perform em-bodied action control, with frequent communication with a goal-conditioned active perception scheme. Specifically, MP5 is developed on top of recent advances in Multimodal Large Language Models (MLLMs), and the system is mod-ulated into functional modules that can be scheduled and collaborated to ultimately solve pre-defined context- and process-dependent tasks. Extensive experiments prove that MP5 can achieve a 22% success rate on difficult process-dependent tasks and a 91 % success rate on tasks that heav-ily depend on the context. Moreover, MP5 exhibits a re-markable ability to address many open-ended tasks that are entirely novel. Please see the project page at https: //iranqin. github.io/MP5. github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen et al.NeurIPS 2024 · 104 citations
- OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following AgentsZihao Wang, Shaofei Cai, Zhancun Mu, Haowei Lin et al.NeurIPS 2024 · 37 citations
- WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM AgentsSiyu Zhou, Tianyi Zhou, Yijun Yang, Guodong Long et al.NeurIPS 2025 · 18 citations
- SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for RoboticsMengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han et al.CVPR 2026 · 10 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
Related papers
- Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task AgentsZihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu et al.NeurIPS 2023 · 178 citations
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng et al.CVPR 2026
- RL-GPT: Integrating Reinforcement Learning and Code-as-policyShaoteng Liu, Haoqi Yuan, Minda Hu, Yanwei Li et al.NeurIPS 2024 · 48 citations
- ADAM: An Embodied Causal Agent in Open-World EnvironmentsShu Yu, Chaochao LuICLR 2025
- LARM: Large Auto-Regressive Model for Long-Horizon Embodied IntelligenceZhuoling Li, Xiaogang Xu, Zhenhua Xu, Ser-Nam Lim et al.ICML 2025
