Monocular Camera-Based Point-Goal Navigation by Learning Depth Channel and Cross-Modality Pyramid Fusion
Tianqi Tang, Heming Du, Xin Yu, Yi Yang
Abstract
For a monocular camera-based navigation system, if we could effectively explore scene geometric cues from RGB images, the geometry information will significantly facilitate the efficiency of the navigation system. Motivated by this, we propose a highly efficient point-goal navigation framework, dubbed Geo-Nav. In a nutshell, our Geo-Nav consists of two parts: a visual perception part and a navigation part. In the visual perception part, we firstly propose a Self-supervised Depth Estimation network (SDE) specially tailored for the monocular camera-based navigation agent. Our SDE learns a mapping from an RGB input image to its corresponding depth image by exploring scene geometric constraints in a self-consistency manner. Then, in order to achieve a representative visual representation from the RGB inputs and learned depth images, we propose a Cross-modality Pyramid Fusion module (CPF). Concretely, our CPF computes a patch-wise cross-modality correlation between different modal features and exploits the correlation to fuse and enhance features at each scale. Thanks to the patch-wise nature of our CPF, we can fuse feature maps at high resolution, allowing our visual network to perceive more image details. In the navigation part, our extracted visual representations are fed to a navigation policy network to learn how to map the visual representations to agent actions effectively. Extensive experiments on a widely-used multiple-room environment Gibson demonstrate that Geo-Nav outperforms the state-of-the-art in terms of efficiency and effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9b9a5cf-deaa-4a0a-8c5d-349174a8ba09Cited by top-tier papers5
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu et al.ICCV 2023 · 136 citations
- FloNa: Floor Plan Guided Embodied Visual NavigationJiaxin Li, Weiqi Huang, Zan Wang, Wei Liang et al.AAAI 2025 · 11 citations
- Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-SourcesZhanbo Shi, Lin Zhang, Linfei Li, Ying ShenAAAI 2025 · 8 citations
- Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion ControlHao Ren, Zetong Bi, Yiming Zeng, Le Zheng et al.ICML 2026 · 1 citation
- Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation StatesHeming Du, Lincheng Li, Zi Huang, Xin YuCVPR 2023
Builds on9
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- SplitNet: Sim2Sim and Task2Task Transfer for Embodied Visual NavigationDaniel Gordon, Abhishek Kadian, Devi Parikh, Judy Hoffman et al.ICCV 2019 · 80 citations
- Situational Fusion of Visual Representation for Visual NavigationWilliam B. Shen, Danfei Xu, Yuke Zhu, Li Fei-Fei et al.ICCV 2019 · 70 citations
- Domain Consensus Clustering for Universal Domain AdaptationGuangrui Li, Guoliang Kang, Yi Zhu, Yunchao Wei et al.CVPR 2021
Related papers
- Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationSixian Zhang, Xinyao Yu, Xinhang Song, Xiaohan Wang et al.CVPR 2024 · 14 citations
- Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationBolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang et al.ACM MM 2024 · 4 citations
- Neural Topological SLAM for Visual NavigationDevendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, Saurabh GuptaCVPR 2020
- MGNet: Monocular Geometric Scene Understanding for Autonomous DrivingMarkus Schön, Michael Buchholz, Klaus DietmayerICCV 2021 · 60 citations
- 3D-Aware Object Goal Navigation via Simultaneous Exploration and IdentificationJiazhao Zhang, Liu Dai, Fanpeng Meng, Qingnan Fan et al.CVPR 2023
