Visual Localization using Imperfect 3D Models from the Internet
Vojtech Panek, Zuzana Kukelova, Torsten Sattler
Abstract
Visual localization is a core component in many applications, including augmented reality (AR). Localization algorithms compute the camera pose of a query image w.r.t. a scene representation, which is typically built from images. This often requires capturing and storing large amounts of data, followed by running Structure-from-Motion (SfM) algorithms. An interesting, and underexplored, source of data for building scene representations are 3D models that are readily available on the Internet, e.g., hand-drawn CAD models, 3D models generated from building footprints, or from aerial images. These models allow to perform visual localization right away without the time-consuming scene capturing and model building steps. Yet, it also comes with challenges as the available 3D models are often imperfect reflections of reality. E.g., the models might only have generic or no textures at all, might only provide a simple approximation of the scene geometry, or might be stretched. This paper studies how the imperfections of these models affect localization accuracy. We create a new benchmark for this task and provide a detailed experimental evaluation based on multiple 3D models per scene. We show that 3D models from the Internet show promise as an easy-to-obtain scene representation. At the same time, there is significant room for improvement for visual localization pipelines. To foster research on this interesting and challenging task, we release our benchmark at v-pnk.github.io/cadloc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01a0850d-eee7-4a3b-a36a-61277a7dca36Cited by top-tier papers11
- LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe AlignmentJuelin Zhu, Shen Yan, Long Wang, Shengyue Zhang et al.NeurIPS 2024 · 17 citations
- The Unreasonable Effectiveness of Pre-Trained Features for Camera Pose RefinementGabriele Trivigno, Carlo Masone, Barbara Caputo, Torsten SattlerCVPR 2024 · 10 citations
- LoD-Loc v2: Aerial Visual Localization Over Low Level-of-Detail City Models using Explicit Silhouette AlignmentJuelin Zhu, Shuaibang Peng, Long Wang, Hanlin Tan et al.ICCV 2025 · 3 citations
- LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette AlignmentShuaibang Peng, Juelin Zhu, Xia Li, Kun Yang et al.CVPR 2026 · 2 citations
- NormalLoc: Visual Localization on Textureless 3D Models using Surface NormalsJiro Abe, Gaku Nakano, Kazumine OguraICCV 2025 · 2 citations
Builds on15
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 163 citations
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi et al.ICCV 2019 · 163 citations
- On the Limits of Pseudo Ground Truth in Visual Camera Re-localisationEric Brachmann, Martin Humenberger, Carsten Rother, Torsten SattlerICCV 2021 · 82 citations
- Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual LocalizationHajime Taira, Ignacio Rocco, Jirí Sedlár, Masatoshi Okutomi et al.ICCV 2019 · 54 citations
Related papers
- CrowdDriven: A New Challenging Dataset for Outdoor Visual LocalizationAra Jafarzadeh, Manuel López-Antequera, Pau Gargallo, Yubin Kuang et al.ICCV 2021 · 18 citations
- CroCoDL: Cross-device Collaborative Dataset for LocalizationHermann Blum, Alessandro Mercurio, Joshua O'Reilly, Tim Engelbracht et al.CVPR 2025
- Large-Scale Localization Datasets in Crowded Indoor SpacesDonghwan Lee, Soo-Hyun Ryu, Suyong Yeon, Yonghan Lee et al.CVPR 2021
- CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic DataQi Yan, Jianhao Zheng, Simon Reding, Shanci Li et al.CVPR 2022 · 23 citations
- SplatLoc: 3D Gaussian Splatting-based Visual Localization for Augmented RealityHongjia Zhai, Xiyu Zhang, Boming Zhao, Hai Li et al.IEEE VR 2025 · 29 citations
