Learning Multi-Scene Absolute Pose Regression with Transformers
Yoli Shavit, Ron Ferens, Yosi Keller
摘要
Absolute camera pose regressors estimate the position and orientation of a camera from the captured image alone. Typically, a convolutional backbone with a multi-layer perceptron head is trained using images and pose labels to embed a single reference scene at a time. Recently, this scheme was extended for learning multiple scenes by replacing the MLP head with a set of fully connected layers. In this work, we propose to learn multi-scene absolute camera pose regression with Transformers, where encoders are used to aggregate activation maps with self-attention and decoders transform latent features and scenes encoding into candidate pose predictions. This mechanism allows our model to focus on general features that are informative for localization while embedding multiple scenes in parallel. We evaluate our method on commonly benchmarked indoor and outdoor datasets and show that it surpasses both multi-scene and state-of-the-art single-scene absolute pose regressors. We make our code publicly available from https://github.com/yolish/multi-scene-pose-transformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Leveraging Equivariant Features for Absolute Pose RegressionMohamed Adel Musallam, Vincent Gaudillière, Miguel Ortiz del Castillo, Kassem Al Ismaeil 等CVPR 2022 · 被引用 27 次
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel 等CVPR 2024 · 被引用 24 次
- NeRF-IBVS: Visual Servo Based on NeRF for Visual Localization and NavigationYuanze Wang, Yichao Yan, Dianxi Shi, Wenhan Zhu 等NeurIPS 2023 · 被引用 20 次
- OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesTao Xie, Kun Dai, Siyi Lu, Ke Wang 等ICCV 2023 · 被引用 16 次
- The Unreasonable Effectiveness of Pre-Trained Features for Camera Pose RefinementGabriele Trivigno, Carlo Masone, Barbara Caputo, Torsten SattlerCVPR 2024 · 被引用 10 次
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao 等AAAI 2020 · 被引用 189 次
- CamNet: Coarse-to-Fine Retrieval for Camera Re-LocalizationMingyu Ding, Zhe Wang, Jiankai Sun, Jianping Shi 等ICCV 2019 · 被引用 163 次
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 被引用 136 次
相关 Paper
- Activating Self-Attention for Multi-Scene Absolute Pose RegressionMiso Lee, Jihwan Kim, Jae-Pil HeoNeurIPS 2024 · 被引用 4 次
- Map-Relative Pose Regression for Visual Re-LocalizationShuai Chen, Tommaso Cavallari, Victor Adrian Prisacariu, Eric BrachmannCVPR 2024
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 被引用 279 次
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han 等CVPR 2022 · 被引用 45 次
- FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationChris Rockwell, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park 等CVPR 2024
