ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
Leonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart, Sowmya Munukutla, Victor Adrian Prisacariu, Eric Brachmann
摘要
Scene coordinate regression (SCR) has established itself as a promising learning-based approach to visual relocalization. After mere minutes of scene-specific training, SCR models estimate camera poses of query images with high accuracy. Still, SCR methods fall short of the generalization capabilities of more classical feature-matching approaches. When imaging conditions of query images, such as lighting or viewpoint, are too different from the training views, SCR models fail. Failing to generalize is an inherent limitation of previous SCR frameworks, since their training objective is to encode the training views in the weights of the coordinate regressor itself. The regressor essentially overfits to the training views, by design. We propose to separate the coordinate regressor and the map representation into a generic transformer and a scene-specific map code. This separation allows us to pre-train the transformer on tens of thousands of scenes. More importantly, it allows us to train the transformer to generalize from mapping images to unseen query images during pre-training. We demonstrate on multiple challenging relocalization datasets that our method, ACE-G, leads to significantly increased robustness while keeping the computational footprint attractive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian SplattingYingdong Gu, Shaocheng Yan, Zhenjun Zhao, Yuan Kou 等CVPR 2026 · 被引用 3 次
- Simple but Effective Triplet-Based Compression Strategies for Compact Visual LocalizationTorsten Sattler, Zuzana KukelovaCVPR 2026
- Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph OptimizationTze Ho Elden Tse, Jizong Peng, Angela YaoCVPR 2026
- Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian SplattingHuaqi Tao, Bingxi Liu, Guangcheng Chen, Fulin Tang 等CVPR 2026
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 被引用 659 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao 等AAAI 2020 · 被引用 189 次
相关 Paper
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel 等CVPR 2024 · 被引用 24 次
- CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual LocalizationXindong Mao, Hang Li, Yuchen Wu, Jiahe Li 等CVPR 2026
- Scene Coordinate Reconstruction PriorsWenjing Bian, Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu 等ICCV 2025 · 被引用 2 次
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 被引用 163 次
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li 等ICCV 2019 · 被引用 105 次
