End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent Phenomenon
Guillaume Bono, Leonid Antsfeld, Boris Chidlovskii, Philippe Weinzaepfel, Christian Wolf
摘要
Most recent work in goal oriented visual navigation resorts to large-scale machine learning in simulated environments. The main challenge lies in learning compact representations generalizable to unseen environments and in learning high-capacity perception modules capable of reasoning on high-dimensional input. The latter is particularly difficult when the goal is not given as a category ("ObjectNav") but as an exemplar image ("ImageNav"), as the perception module needs to learn a comparison strategy requiring to solve an underlying visual correspondence problem. This has been shown to be difficult from reward alone or with standard auxiliary tasks. We address this problem through a sequence of two pretext tasks, which serve as a prior for what we argue is one of the main bottleneck in perception, extremely wide-baseline relative pose estimation and visibility prediction in complex scenes. The first pretext task, cross-view completion is a proxy for the underlying visual correspondence problem, while the second task addresses goal detection and finding directly. We propose a new dual encoder with a large-capacity binocular ViT model and show that correspondence solutions naturally emerge from the training signals. Experiments show significant improvements and SOTA performance on the two benchmarks, ImageNav and the Instance-ImageNav variant, where camera intrinsics and height differ between observation and goal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Alligat0R: Pre-Training through Covisibility Segmentation for Relative Camera Pose RegressionThibaut Loiseau, Guillaume Bourmaud, Vincent LepetitNeurIPS 2025 · 被引用 11 次
- Kinaema: a recurrent sequence model for memory and pose in motionMert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci 等NeurIPS 2025 · 被引用 3 次
- EmbodiedSplat: Personalized Real-To-Sim-To-Real Navigation with Gaussian Splats From a Mobile DeviceGunjan Chhablani, Xiaomeng Ye, Muhammad Zubair Irshad, Zsolt KiraICCV 2025 · 被引用 2 次
- Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems ApproachSteeven Janny, Hervé Poirier, Leonid Antsfeld, Guillaume Bono 等CVPR 2025
- GOAT-Bench: A Benchmark for Multi-Modal Lifelong NavigationMukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra 等CVPR 2024
它引用的顶会 Paper23
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
相关 Paper
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier 等NeurIPS 2022 · 被引用 189 次
- Navigating to Objects Specified by ImagesJacob Krantz, Théophile Gervet, Karmesh Yadav, Austin S. Wang 等ICCV 2023 · 被引用 70 次
- ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene ImaginationXinxin Zhao, Wenzhe Cai, Likun Tang, Teng WangICLR 2025
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationBadi Li, Renjie Lu, Yu Zhou, Jingke Meng 等NeurIPS 2025 · 被引用 5 次
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman 等NeurIPS 2022 · 被引用 344 次
