TransGeo: Transformer Is All You Need for Cross-view Image Geo-localization
Sijie Zhu, Mubarak Shah, Chen Chen
Abstract
The dominant CNN-based methods for cross-view image geo-localization rely on polar transform and fail to model global correlation. We propose a pure transformer-based approach (TransGeo) to address these limitations from a different perspective. TransGeo takes full advantage of the strengths of transformer related to global information modeling and explicit position information encoding. We further leverage the flexibility of transformer input and propose an attention-guided non-uniform cropping method, so that uninformative image patches are removed with negligible drop on performance to reduce computation cost. The saved computation can be reallocated to increase resolution only for informative patches, resulting in performance improvement with no additional computation cost. This "attend and zoom-in" strategy is highly similar to human behavior when observing images. Remarkably, TransGeo achieves stateof-the-art results on both urban and rural datasets, with significantly less computation cost than CNN-based methods. It does not rely on polar transform and infers faster than CNN-based methods. Code is available at https: //github.com/Jeff-Zilence/TransGeo2022 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1a4be4a-2b47-4f8b-bdb2-f7b0ae0aa4f9Cited by top-tier papers42
- Sample4Geo: Hard Negative Sampling For Cross-View Geo-LocalisationFabian Deuser, Konrad Habel, Norbert OswaldICCV 2023 · 161 citations
- Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View TransformerYujiao Shi, Fei Wu, Akhil Perincherry, Ankit Vora et al.ICCV 2023 · 60 citations
- Game4Loc: A UAV Geo-Localization Benchmark from Game DataYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuAAAI 2025 · 35 citations
- Beyond Geo-localization: Fine-grained Orientation of Street-view Images by Cross-view Matching with Satellite ImageryWenmiao Hu, Yichen Zhang, Yuxuan Liang, Yifang Yin et al.ACM MM 2022 · 31 citations
- Aligning Geometric Spatial Layout in Cross-View Geo-Localization via Feature RecombinationQingwang Zhang, Yingying ZhuAAAI 2024 · 28 citations
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
- Bridging the Domain Gap for Ground-to-Aerial Image MatchingKrishna Regmi, Mubarak ShahICCV 2019 · 191 citations
Related papers
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper et al.IEEE VR 2023 · 22 citations
- Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and ScenesBrandon Clark, Alec Kerrigan, Parth Parag Kulkarni, Vicente Vivanco Cepeda et al.CVPR 2023
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 115 citations
- Breaking Rectangular Shackles: Cross-View Object Segmentation for Fine-Grained Object Geo-LocalizationQingwang Zhang, Yingying ZhuICCV 2025 · 2 citations
- Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View MatchingYujiao Shi, Xin Yu, Dylan Campbell, Hongdong LiCVPR 2020
