RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization
Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar
Abstract
We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (rendered as depth images). Prior works were demonstrated on small datasets and did not lend themselves to scaling up for large-scale applications. To enable large-scale evaluation, we introduce a new dataset containing over 550K pairs (covering 143 km2 area) of RGB and aerial LIDAR depth images. We propose a novel joint embedding based method that effectively combines the appearance and semantic cues from both modalities to handle drastic cross-modal variations. Experiments on the proposed dataset show that our model achieves a strong result of a median rank of 5 in matching across a large test set of 50K location pairs collected from a 14km^2 area. This represents a significant advancement over prior works in performance and scale. We conclude with qualitative results to highlight the challenging nature of this task and the benefits of the proposed model. Our work provides a foundation for further research in cross-modal visual localization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.ICCV 2023 · 69 citations
- Cross-View Visual Geo-Localization for Outdoor Augmented RealityNiluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper et al.IEEE VR 2023 · 22 citations
Builds on2
Related papers
- MMGeo: Multimodal Compositional Geo-Localization for UAVsYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuICCV 2025 · 5 citations
- LCD: Learned Cross-Domain Descriptors for 2D-3D MatchingQuang-Hieu Pham, Mikaela Angelina Uy, Binh-Son Hua, Duc Thanh Nguyen et al.AAAI 2020 · 94 citations
- L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing ImageryZiwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia et al.NeurIPS 2025 · 3 citations
- CrossLoc3D: Aerial-Ground Cross-Source 3D Place RecognitionTianrui Guan, Aswath Muthuselvam, Montana Hoover, Xijun Wang et al.ICCV 2023 · 24 citations
- AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional RelationsJunli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang et al.ICCV 2025 · 5 citations
