CrossLoc: Scalable Aerial Localization Assisted by Multimodal Synthetic Data
Qi Yan, Jianhao Zheng, Simon Reding, Shanci Li, Iordan Doytchinov
摘要
We present a visual localization system that learns to estimate camera poses in the real world with the help of synthetic data. Despite significant progress in recent years, most learning-based approaches to visual localization target at a single domain and require a dense database of geo-tagged images to function well. To mitigate the data scarcity issue and improve the scalability of the neural localization models, we introduce TOPO-DataGen, a versatile synthetic data generation tool that traverses smoothly between the real and virtual world, hinged on the geographic camera viewpoint. New large-scale sim-to-real benchmark datasets are proposed to showcase and evaluate the utility of the said synthetic data. Our experiments reveal that synthetic data generically enhances the neural network performance on real data. Furthermore, we introduce CrossLoc, a cross-modal visual representation learning approach to pose estimation that makes full use of the scene coordinate ground truth via self-supervision. Without any extra data, CrossLoc significantly outperforms the state-of-the-art methods and achieves substantially higher real-data sample efficiency. Our code and datasets are all available at crossloc. github. io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe AlignmentJuelin Zhu, Shen Yan, Long Wang, Shengyue Zhang 等NeurIPS 2024 · 被引用 17 次
- Self-Prompting Analogical Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Song Tang 等AAAI 2025 · 被引用 10 次
- UAVScenes: A Multi-Modal Dataset for UAVsSijie Wang, Siqi Li, Yawei Zhang, Shangshu Yu 等ICCV 2025 · 被引用 9 次
- OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial PerspectiveMarkus Gross, Sai B. Matha, Aya Fahmy, Rui Song 等CVPR 2026 · 被引用 7 次
- MMGeo: Multimodal Compositional Geo-Localization for UAVsYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuICCV 2025 · 被引用 5 次
它引用的顶会 Paper12
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas 等ICML 2020 · 被引用 651 次
- Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D ScansAinaz Eftekhar, Alexander Sax, Jitendra Malik, Amir ZamirICCV 2021 · 被引用 422 次
- Neural-Guided RANSAC: Learning Where to Sample Model HypothesesEric Brachmann, Carsten RotherICCV 2019 · 被引用 282 次
- AtLoc: Attention Guided Camera LocalizationBing Wang, Changhao Chen, Chris Xiaoxuan Lu, Peijun Zhao 等AAAI 2020 · 被引用 189 次
- Expert Sample Consensus Applied to Camera Re-LocalizationEric Brachmann, Carsten RotherICCV 2019 · 被引用 136 次
相关 Paper
- 360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-Device QueriesHuajian Huang, Changkun Liu, Yipeng Zhu, Hui Cheng 等CVPR 2024
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- PNeRFLoc: Visual Localization with Point-Based Neural Radiance FieldsBoming Zhao, Luwei Yang, Mao Mao, Hujun Bao 等AAAI 2024 · 被引用 31 次
- Rethinking Visual Geo-localization for Large-Scale ApplicationsGabriele Moreno Berton, Carlo Masone, Barbara CaputoCVPR 2022 · 被引用 235 次
- Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationSiyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai 等CVPR 2025
