BEAR: Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis
Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma, Ziyan Chen, Yaokun Han, Renrui Zhang, Zitiantao Lin, Yizhe Zhu, Shiji Xin, Yijian Huang, Boce Hu
摘要
Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address this, we introduce BEAR, where we divide embodied tasks into 14 atomic skills for skill-level evaluation. BEAR comprises 4,469 interleaved image–video–text entries across 14 skills in 6 categories, ranging from low-level perception to high-level planning. We evaluate 20 MLLMs on BEAR under a hierarchical skill-level diagnosis framework and discover that (1) perceptual capabilities are major bottlenecks behind reasoning failures, and (2) models fail due to unstable spatiotemporal modeling which remain unexposed in previous benchmarks. Furthermore, building on these insights, we propose BEAR-Agent, a multimodal conversable agent that augments MLLMs with visual and spatial tools. It substantially enhances MLLMs’ performance across skills, yielding a relative improvement of 17.5% on GPT-5 on BEAR and outperform baselines by a large margin in both simulation and real-robot experiments across models. We provide our project website at https://bear-official66.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 被引用 659 次
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li 等ICML 2024 · 被引用 184 次
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Mutimodal ModelsXingrui Wang, Wufei Ma, Tiezheng Zhang, Celso M. de Melo 等CVPR 2025
相关 Paper
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao 等ICML 2025
- ERGeoBench: A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language ModelsKaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou 等ICML 2026
- VSP: Diagnosing the Dual Challenges of Perception and Reasoning in Spatial Planning Tasks for MLLMSQiucheng Wu, Handong Zhao, Michael Saxon, Trung Bui 等ICCV 2025 · 被引用 1 次
- How Foundational Skills Influence VLM-based Embodied Agents: A Native PerspectiveBo Peng, Pi Bu, Keyu Pan, Xinrun Xu 等AAAI 2026 · 被引用 1 次
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan 等CVPR 2026 · 被引用 2 次
