Augmented Commonsense Knowledge for Remote Object Grounding
Bahram Mohammadi, Yicong Hong, Yuankai Qi, Qi Wu, Shirui Pan, Javen Qinfeng Shi
Abstract
The vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environments. Most of the existing methods employ the entire image or object features to represent navigable viewpoints. However, these representations are insufficient for proper action prediction, especially for the REVERIE task, which uses concise high-level instructions, such as “Bring me the blue cushion in the master bedroom”. To address enhancing representation, we propose an augmented commonsense knowledge model (ACK) to leverage commonsense information as a spatio-temporal knowledge graph for improving agent navigation. Specifically, the proposed approach involves constructing a knowledge base by retrieving commonsense information from ConceptNet, followed by a refinement module to remove noisy and irrelevant knowledge. We further present ACK which consists of knowledge graph-aware cross-modal and concept aggregation modules to enhance visual representation and visual-textual data alignment by integrating visible objects, commonsense knowledge, and concept history, which includes object and knowledge temporal information. Moreover, we add a new pipeline for the commonsense-based decision-making process which leads to more accurate local action prediction. Experimental results demonstrate our proposed model noticeably outperforms the baseline and archives the state-of-the-art on the REVERIE benchmark. The source code is available at https://github.com/Bahram-Mohammadi/ACK.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dfdd614-7532-4a62-ba54-7d28e00c1dd7Cited by top-tier papers3
- COSMO: Combination of Selective Memorization for Low-Cost Vision-and-Language NavigationSiqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan et al.ICCV 2025 · 3 citations
- NavQ: Learning a Q-Model for Foresighted Vision-and-Language NavigationPeiran Xu, Xicheng Gong, Yadong MuICCV 2025 · 2 citations
- Run, Ruminate, and Regulate: A Dual-process Thinking System for Vision-and-Language NavigationYu Zhong, Zihao Zhang, Rui Zhang, Lingdong Huang et al.AAAI 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 228 citations
- Airbert: In-domain Pretraining for Vision-and-Language NavigationPierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev et al.ICCV 2021 · 185 citations
- Think Global, Act Local: Dual-scale Graph Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.CVPR 2022 · 150 citations
Related papers
- Room-and-Object Aware Knowledge Reasoning for Remote Embodied Referring ExpressionChen Gao, Jinyu Chen, Si Liu, Luting Wang et al.CVPR 2021
- KERM: Knowledge Enhanced Reasoning for Vision-and-Language NavigationXiangyang Li, Zihan Wang, Jiahao Yang, Yaowei Wang et al.CVPR 2023
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language NavigationZihan Wang, Seungjun Lee, Gim Hee LeeNeurIPS 2025 · 36 citations
- Cross-modal Map Learning for Vision and Language NavigationGeorgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan et al.CVPR 2022 · 2 citations
- Layout-Aware Dreamer for Embodied Visual Referring Expression GroundingMingxiao Li, Zehao Wang, Tinne Tuytelaars, Marie-Francine MoensAAAI 2023 · 22 citations
