Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual Navigation
Ningnan Wang, Weihuang Chen, Liming Chen, Haoxuan Ji, Zhongyu Guo, Xuchong Zhang, Hongbin Sun
Abstract
Embodied visual navigation remains a challenging task, as agents must explore unknown environments with limited knowledge. Existing zero-shot studies have shown that incorporating memory mechanisms to support goal-directed behavior can improve long-horizon planning performance. However, they overlook visual frontier boundaries, which fundamentally dictate future trajectories and observations, and fall short of inferring the relationship between partial visual observations and navigation goals. In this paper, we propose Semantic Cognition Over Potential-based Exploration (SCOPE), a zero-shot framework that explicitly leverages frontier information to drive potential-based exploration, enabling more informed and goal-relevant decisions. SCOPE estimates exploration potential with a Vision-Language Model and organizes it into a spatio-temporal potential graph, capturing boundary dynamics to support long-horizon planning. In addition, SCOPE incorporates a self-reconsideration mechanism that revisits and refines prior decisions, enhancing reliability and reducing overconfident errors. Experimental results on two diverse embodied navigation tasks show that SCOPE outperforms state-of-the-art baselines by 4.6% in accuracy. Further analysis demonstrates that its core components lead to improved calibration, stronger generalization, and higher decision quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee173117-d0d5-41ef-8e07-3e83dd836313Builds on28
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou et al.NeurIPS 2024 · 215 citations
- OpenEQA: Embodied Question Answering in the Era of Foundation ModelsArjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta et al.CVPR 2024 · 44 citations
- Trajectory Diffusion for ObjectGoal NavigationXinyao Yu, Sixian Zhang, Xinhang Song, Xiaorong Qin et al.NeurIPS 2024 · 32 citations
- Volumetric Environment Representation for Vision-Language NavigationRui Liu, Wenguan Wang, Yi YangCVPR 2024 · 25 citations
- REGNav: Room Expert Guided Image-Goal NavigationPengna Li, Kangyi Wu, Jingwen Fu, Sanping ZhouAAAI 2025 · 15 citations
Related papers
- UniGoal: Towards Universal Zero-shot Goal-oriented NavigationHang Yin, Xiuwei Xu, Linqing Zhao, Ziwei Wang et al.CVPR 2025
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied NavigationXun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu et al.CVPR 2026 · 19 citations
- STRIDER: Navigation via Instruction-Aligned Structural Decision Space OptimizationDiqi He, Xuehao Gao, Hao Li, Junwei Han et al.NeurIPS 2025 · 8 citations
- BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object NavigationZibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li et al.NeurIPS 2025 · 31 citations
- SCOPE: Evolving Symbolic World for Planning in Open-Ended EnvironmentsYundaichuan Zhan, Minghe Gao, Zhongqi Yue, Wendong Bu et al.ICML 2026
