MCU: An Evaluation Framework for Open-Ended Game Agents
Xinyue Zheng, Haowei Lin, Kaichen He, Zihao Wang, Qiang Fu, Haobo Fu, Zilong Zheng, Yitao Liang
摘要
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce Minecraft Universe (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu 等ICCV 2025 · 被引用 7 次
- Experience Transfer for Multimodal LLM Agents in Minecraft GameChenghao Li, Jun Liu, Songbo Zhang, Huadong Jian 等CVPR 2026 · 被引用 4 次
- World2Minecraft: Occupancy-Driven Simulated Scenes ConstructionLechao Zhang, Haoran Xu, Jingyu Gong, Xuhong Wang 等ICLR 2026 · 被引用 1 次
- Experience-based Knowledge Correction for Robust Planning in MinecraftSeungjoon Lee, Suhwan Kim, Minhyeon Oh, Youngsik Yoon 等ICLR 2026 · 被引用 1 次
- UniCode: Augmenting Evaluation for Code ReasoningXinyue Zheng, Haowei Lin, Shaofei Cai, Yaodong Yang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga 等NeurIPS 2022 · 被引用 458 次
- Benchmarking the Spectrum of Agent CapabilitiesDanijar HafnerICLR 2022 · 被引用 193 次
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba 等NeurIPS 2023 · 被引用 123 次
相关 Paper
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 被引用 121 次
- AgentStudio: A Toolkit for Building General Virtual AgentsLongtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang 等ICLR 2025
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsQuyu Kong, Xu Zhang, Zhenyu Yang, Nolan Gao 等ACL 2026 · 被引用 38 次
- MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated ToolsZikang Guo, Benfeng Xu, Chiwei Zhu, Wentao Hong 等AAAI 2026 · 被引用 19 次
- CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended WorldZoya Volovikova, Gregory Gorbov, Petr Kuderov, Aleksandr Panov 等ACL 2025 · 被引用 2 次
