VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation
Juhye Park, Wooju Lee, Dasol Hong, Changki Sung, Youngwoo Seo, Dongwan Kang, Hyun Myung
Abstract
Accurate global localization is critical for autonomous driving and robotics, but GNSS-based approaches often degrade due to occlusion and multipath effects. As an emerging alternative, cross-view pose estimation predicts the 3-DoF camera pose corresponding to a ground-view image with respect to a geo-referenced satellite image. However, existing methods struggle to bridge the significant viewpoint gap between the ground and satellite views mainly due to limited spatial correspondences. We propose a novel cross-view pose estimation method that constructs view-invariant representations through dual-axis transformation (VIRD). VIRD first applies a polar transformation to the satellite view to facilitate horizontal correspondence, then uses context-enhanced positional attention on the ground and polar-transformed satellite features to mitigate vertical misalignment, explicitly bridging the viewpoint gap. To further strengthen view invariance, we introduce a view-reconstruction loss that encourages the derived representations to reconstruct the original and cross-view images. Experiments on the KITTI and VIGOR datasets demonstrate that VIRD outperforms the state-of-the-art methods without orientation priors, reducing median position and orientation errors by 50.7% and 76.5% on KITTI, and 18.0% and 46.8% on VIGOR, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de49eabe-f93d-47b2-9fb2-aba190a4a4caBuilds on13
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
- Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography EstimatorXiaolong Wang, Runsen Xu, Zhuofan Cui, Zeyu Wan et al.NeurIPS 2023 · 96 citations
- Beyond Cross-view Image Retrieval: Highly Accurate Vehicle Localization Using Satellite ImageYujiao Shi, Hongdong LiCVPR 2022 · 81 citations
- Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View TransformerYujiao Shi, Fei Wu, Akhil Perincherry, Ankit Vora et al.ICCV 2023 · 60 citations
- Learning Dense Flow Field for Highly-accurate Cross-view Camera LocalizationZhenbo Song, Xianghui Ze, Jianfeng Lu, Yujiao ShiNeurIPS 2023 · 37 citations
Related papers
- BevSplat: Resolving Height Ambiguity via Feature-Based Gaussian Primitives for Weakly-Supervised Cross-View LocalizationQiwei Wang, Shaoxun Wu, Yujiao ShiNeurIPS 2025 · 10 citations
- PIDLoc: Cross-View Pose Optimization Network Inspired by PID ControllersWooju Lee, Juhye Park, Dasol Hong, Changki Sung et al.CVPR 2025
- Uncertainty-Aware Vision-Based Metric Cross-View GeolocalizationFlorian Fervers, Sebastian Bullinger, Christoph Bodensteiner, Michael Arens et al.CVPR 2023
- Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View MatchingYujiao Shi, Xin Yu, Dylan Campbell, Hongdong LiCVPR 2020
- FG^2: Fine-Grained Cross-View Localization by Fine-Grained Feature MatchingZimin Xia, Alexandre AlahiCVPR 2025
