MCU: An Evaluation Framework for Open-Ended Game Agents
Xinyue Zheng, Haowei Lin, Kaichen He, Zihao Wang, Qiang Fu, Haobo Fu, Zilong Zheng, Yitao Liang
Abstract
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce Minecraft Universe (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d77a74e7-c6e9-45b1-9b68-ca6d8b650c67Cited by top-tier papers5
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu et al.ICCV 2025 · 7 citations
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian et al.CVPR 2026 · 4 citations
- World2Minecraft: Occupancy-Driven Simulated Scenes ConstructionLechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang et al.ICLR 2026 · 1 citation
- Experience-based Knowledge Correction for Robust Planning in MinecraftSeungjoon Lee, Suhwan Kim, Minhyeon Oh, Youngsik Yoon et al.ICLR 2026 · 1 citation
- UniCode: Augmenting Evaluation for Code ReasoningXinyue Zheng, Haowei Lin, Shaofei Cai, Yaodong Yang et al.ICML 2026 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 193 citations
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba et al.NeurIPS 2023 · 123 citations
Related papers
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
- AgentStudio: A Toolkit for Building General Virtual AgentsLongtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang et al.ICLR 2025
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsQuyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao et al.ACL 2026 · 38 citations
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong et al.AAAI 2026 · 19 citations
- CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended WorldZoya Volovikova, Gregory Gorbov, Petr Kuderov, Aleksandr Panov et al.ACL 2025 · 2 citations
