NavBench: Probing Multimodal Large Language Models for Embodied Navigation
Yanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An, Siqi Zhang, Yutong Xie, Xinyu Wang, Qi Wu
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3,200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge.
Recent work has begun to explore MLLMs' potential in embodied tasks by evaluating their spatial reasoning in 3D environments [7,8]. However, these tasks primarily focus on perception and passive scene understanding, without assessing the model's ability to make decisions or take actions. In comparison, navigation is a core embodied task that involves interpreting natural language instructions, analyzing visual observations, and making a sequence of decisions to reach a goal. Although navigation plays a crucial role in real-world applications, it remains relatively underexplored in the context of MLLMs. Traditional embodied navigation benchmarks, such as Room-to-Room (R2R) [9] ˚Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ffd3642-aeff-4970-b002-8a52e69077f0Cited by top-tier papers8
- pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial ReasoningZhanpeng Luo, Ce Zhang, Silong Yong, Cunxi Dai et al.ICLR 2026 · 15 citations
- Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied ExplorationSen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma et al.CVPR 2026 · 14 citations
- CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor NavigationXia Su, Ruiqi Chen, Benlin Liu, Jingwei Ma et al.CVPR 2026 · 8 citations
- VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation AgentsXunyi Zhao, Gengze Zhou, Qi WuACL 2026 · 3 citations
- From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level GroundingYuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu et al.CVPR 2026 · 2 citations
Builds on31
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
Related papers
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied AgentsRui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao et al.ICML 2025
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li et al.ICML 2024 · 184 citations
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?Yun Li, Yiming Zhang, Tao Lin, XiangRui Liu et al.ICCV 2025 · 4 citations
- How Foundational Skills Influence VLM-based Embodied Agents: A Native PerspectiveBo Peng, Pi Bu, Keyu Pan, Xinrun Xu et al.AAAI 2026 · 1 citation
