SNAP: Self-Supervised Neural Maps for Visual Positioning and Semantic Understanding
Paul-Edouard Sarlin, Eduard Trulls, Marc Pollefeys, Jan Hosang, Simon Lynen
Abstract
Semantic 2D maps are commonly used by humans and machines for navigation purposes, whether it's walking or driving. However, these maps have limitations: they lack detail, often contain inaccuracies, and are difficult to create and maintain, especially in an automated fashion. Can we use raw imagery to automatically create better maps that can be easily interpreted by both humans and machines? We introduce SNAP, a deep network that learns rich neural 2D maps from ground-level and overhead images. We train our model to align neural maps estimated from different inputs, supervised only with camera poses over tens of millions of StreetView images. SNAP can resolve the location of challenging image queries beyond the reach of traditional methods, outperforming the state of the art in localization by a large margin. Moreover, our neural maps encode not only geometry and appearance but also high-level semantics, discovered without explicit supervision. This enables effective pre-training for data-efficient semantic scene understanding, with the potential to unlock cost-efficient creation of more detailed maps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ab88333-35db-42f2-afa9-403d2a767001Cited by top-tier papers8
- LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe AlignmentJuelin Zhu, Shen Yan, Long Wang, Shengyue Zhang et al.NeurIPS 2024 · 17 citations
- Scaling Image Geo-Localization to Continent LevelPhilipp Lindenberger, Paul-Edouard Sarlin, Jan Hosang, Marc Pollefeys et al.NeurIPS 2025 · 11 citations
- BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View LocalizationQiwei Wang, Shaoxun Wu, Yujiao ShiNeurIPS 2025 · 10 citations
- GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View LocalizationShaowen Tong, Zimin Xia, Alexandre Alahi, Xuming He et al.ICCV 2025 · 3 citations
- LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette AlignmentShuaibang Peng, Juelin Zhu, Xia Li, Kun Yang et al.CVPR 2026 · 2 citations
Builds on37
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
Related papers
- OrienterNet: Visual Localization in 2D Public Maps with Neural MatchingPaul-Edouard Sarlin, Daniel DeTone, Tsun-Yi Yang, Armen Avetisyan et al.CVPR 2023
- Predicting Semantic Map Representations From Images Using Pyramid Occupancy NetworksThomas Roddick, Roberto CipollaCVPR 2020
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationJungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto et al.ICCV 2025 · 8 citations
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li et al.ICCV 2019 · 105 citations
- NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera LocalizationShitao Tang, Sicong Tang, Andrea Tagliasacchi, Ping Tan et al.CVPR 2023
