Lune

NeurIPS2025Top-tier venue

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Yanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An, Siqi Zhang, Yutong Xie, Xinyu Wang, Qi Wu

2025Year
27Citations
8Top-tier citations

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3,200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge.

Recent work has begun to explore MLLMs' potential in embodied tasks by evaluating their spatial reasoning in 3D environments [7,8]. However, these tasks primarily focus on perception and passive scene understanding, without assessing the model's ability to make decisions or take actions. In comparison, navigation is a core embodied task that involves interpreting natural language instructions, analyzing visual observations, and making a sequence of decisions to reach a goal. Although navigation plays a crucial role in real-world applications, it remains relatively underexplored in the context of MLLMs. Traditional embodied navigation benchmarks, such as Room-to-Room (R2R) [9] ˚Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7ffd3642-aeff-4970-b002-8a52e69077f0

Cited by top-tier papers8

Ask how each one uses it

Builds on31

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines