ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
Leonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart, Sowmya Munukutla, Victor Adrian Prisacariu, Eric Brachmann
Abstract
Scene coordinate regression (SCR) has established itself as a promising learning-based approach to visual relocalization. After mere minutes of scene-specific training, SCR models estimate camera poses of query images with high accuracy. Still, SCR methods fall short of the generalization capabilities of more classical feature-matching approaches. When imaging conditions of query images, such as lighting or viewpoint, are too different from the training views, SCR models fail. Failing to generalize is an inherent limitation of previous SCR frameworks, since their training objective is to encode the training views in the weights of the coordinate regressor itself. The regressor essentially overfits to the training views, by design. We propose to separate the coordinate regressor and the map representation into a generic transformer and a scene-specific map code. This separation allows us to pre-train the transformer on tens of thousands of scenes. More importantly, it allows us to train the transformer to generalize from mapping images to unseen query images during pre-training. We demonstrate on multiple challenging relocalization datasets that our method, ACE-G, leads to significantly increased robustness while keeping the computational footprint attractive.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian SplattingYingdong Gu, Shaocheng Yan, Zhenjun Zhao, Yuan Kou et al.CVPR 2026 · 3 citations
- Simple but Effective Triplet-Based Compression Strategies for Compact Visual LocalizationTorsten Sattler, Zuzana KukelovaCVPR 2026
- Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph OptimizationTze Ho Elden Tse, Jizong Peng, Angela YaoCVPR 2026
- Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian SplattingHuaqi Tao, Bingxi Liu, Guangcheng Chen, Fulin Tang et al.CVPR 2026
Builds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao et al.AAAI 2020 · 189 citations
Related papers
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel et al.CVPR 2024 · 24 citations
- CoLoR: The Devil is in Scene Coordinate Regression for Large-Scale Visual LocalizationXindong Mao, Hang Li, Yuchen Wu, Jiahe Li et al.CVPR 2026
- Scene Coordinate Reconstruction PriorsWenjing Bian, Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu et al.ICCV 2025 · 2 citations
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 163 citations
- SANet: Scene Agnostic Network for Camera LocalizationLuwei Yang, Ziqian Bai, Chengzhou Tang, Honghua Li et al.ICCV 2019 · 105 citations
