Semantic-Guided Camera Ray Regression for Visual Localization
Yesheng Zhang, Xu Zhao
Abstract
This work presents a novel framework for Visual Localization (VL), that is, regressing camera rays from query images to derive camera poses. As an overparameterized representation of the camera pose, camera rays possess superior robustness in optimization. Of particular importance, Camera Ray Regression (CRR) is privacy-preserving, rendering it a viable VL approach for real-world applications. Thus, we introduce DINO-based Multi-Mappers, coined DIMM, to achieve VL by CRR. DIMM utilizes DINO as a sceneagnostic encoder to obtain powerful features from images. To mitigate ambiguity, the features integrate both local and global perception, as well as potential geometric constraint. Then, a scene-specific mapper head regresses camera rays from these features. It incorporates a semantic attention module for soft fusion of multiple mappers, utilizing the rich semantic information in DINO features. In extensive experiments on both indoor and outdoor datasets, our methods showcase impressive performance, revealing a promising direction for advancements in VL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c8a7c5d-8a04-4852-9f84-24693254852cBuilds on19
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 163 citations
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 136 citations
Related papers
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
- Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual LocalizationMaxime Pietrantoni, Gabriela Csurka, Torsten SattlerCVPR 2025
- On the Generalization Capacities of MLLMs for Spatial IntelligenceGongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang et al.ICLR 2026 · 10 citations
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-TrainingLeonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart et al.ICCV 2025 · 5 citations
- GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian SplattingChangkun Liu, Shuai Chen, Yash Bhalgat, Siyan Hu et al.ICLR 2025 · 1 citation
