Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation
Sixian Zhang, Xinyao Yu, Xinhang Song, Xiaohan Wang, Shuqiang Jiang
摘要
The Object Goal navigation (ObjectNav) task requires the agent to navigate to a specified target in an unseen environment. Since the environment layout is unknown, the agent needs to infer the unknown contextual objects from partially observations, thereby deducing the likely location of the target. Previous end-to-end RL methods capture contextual relationships through implicit representations while they lack notion of geometry. Alternatively, modular methods construct local maps for recording the observed geometric structure of unseen environment, however, lacking the reasoning of contextual relation limits the exploration efficiency. In this work, we propose the self-supervised generative map (SGM), a modular method that learns the explicit context relation via self-supervised learning. The SGM is trained to leverage both episodic observations and general knowledge to reconstruct the masked pixels of a cropped global map. During navigation, the agent maintains an incomplete local semantic map, meanwhile, the unknown regions of the local map are generated by the pretrained SGM. Based on the generated map, the agent sets the predicted location of the target as the goal and moves towards it. Experiments on Gibson, MP3D and HM3D show the effectiveness of our method. The code is available at https://github.com/sx-zhang/SGM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Trajectory Diffusion for ObjectGoal NavigationXinyao Yu, Sixian Zhang, Xinhang Song, Xiaorong Qin 等NeurIPS 2024 · 被引用 32 次
- BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object NavigationZibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li 等NeurIPS 2025 · 被引用 31 次
- Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language NavigationPingrui Zhang, Yifei Su, Pengyuan Wu, Dong An 等CVPR 2026 · 被引用 19 次
- REGNav: Room Expert Guided Image-Goal NavigationPengna Li, Kangyi Wu, Jingwen Fu, Sanping ZhouAAAI 2025 · 被引用 15 次
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalZebin Yang, Sunjian Zheng, Tong Xie, Tianshi Xu 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 被引用 690 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
相关 Paper
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationBadi Li, Renjie Lu, Yu Zhou, Jingke Meng 等NeurIPS 2025 · 被引用 5 次
- PEANUT: Predicting and Navigating to Unseen TargetsAlbert J. Zhai, Shenlong WangICCV 2023 · 被引用 52 次
- Learning to Map for Active Semantic Goal NavigationGeorgios Georgakis, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh 等ICLR 2022 · 被引用 105 次
- GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceShuaihang Yuan, Hao Huang, Yu Hao, Congcong Wen 等NeurIPS 2024 · 被引用 42 次
- Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationBolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang 等ACM MM 2024 · 被引用 4 次
