Semantic-Guided Camera Ray Regression for Visual Localization
Yesheng Zhang, Xu Zhao
摘要
This work presents a novel framework for Visual Localization (VL), that is, regressing camera rays from query images to derive camera poses. As an overparameterized representation of the camera pose, camera rays possess superior robustness in optimization. Of particular importance, Camera Ray Regression (CRR) is privacy-preserving, rendering it a viable VL approach for real-world applications. Thus, we introduce DINO-based Multi-Mappers, coined DIMM, to achieve VL by CRR. DIMM utilizes DINO as a sceneagnostic encoder to obtain powerful features from images. To mitigate ambiguity, the features integrate both local and global perception, as well as potential geometric constraint. Then, a scene-specific mapper head regresses camera rays from these features. It incorporates a semantic attention module for soft fusion of multiple mappers, utilizing the rich semantic information in DINO features. In extensive experiments on both indoor and outdoor datasets, our methods showcase impressive performance, revealing a promising direction for advancements in VL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 被引用 163 次
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 被引用 136 次
相关 Paper
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu 等CVPR 2026 · 被引用 23 次
- Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual LocalizationMaxime Pietrantoni, Gabriela Csurka, Torsten SattlerCVPR 2025
- On the Generalization Capacities of MLLMs for Spatial IntelligenceGongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang 等ICLR 2026 · 被引用 10 次
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-TrainingLeonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart 等ICCV 2025 · 被引用 5 次
- GS-CPR: Efficient Camera Pose Refinement via 3D Gaussian SplattingChangkun Liu, Shuai Chen, Yash Bhalgat, Siyan Hu 等ICLR 2025 · 被引用 1 次
