DistilVPR: Cross-Modal Knowledge Distillation for Visual Place Recognition
Sijie Wang, Rui She, Qiyu Kang, Xingchao Jian, Kai Zhao, Yang Song, Wee Peng Tay
Abstract
The utilization of multi-modal sensor data in visual place recognition (VPR) has demonstrated enhanced performance compared to single-modal counterparts. Nonetheless, integrating additional sensors comes with elevated costs and may not be feasible for systems that demand lightweight operation, thereby impacting the practical deployment of VPR. To address this issue, we resort to knowledge distillation, which empowers single-modal students to learn from cross-modal teachers without introducing additional sensors during inference. Despite the notable advancements achieved by current distillation approaches, the exploration of feature relationships remains an under-explored area. In order to tackle the challenge of cross-modal distillation in VPR, we present Dis-tilVPR, a novel distillation pipeline for VPR. We propose leveraging feature relationships from multiple agents, including self-agents and cross-agents for teacher and student neural networks. Furthermore, we integrate various manifolds, characterized by different space curvatures for exploring feature relationships. This approach enhances the diversity of feature relationships, including Euclidean, spherical, and hyperbolic relationship modules, thereby enhancing the overall representational capacity. The experiments demonstrate that our proposed pipeline achieves state-of-the-art performance compared to other distillation baselines. We also conduct necessary ablation studies to show design effectiveness. The code is released at: https://github.com/sijieaaa/DistilVPR
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ded9824c-c07a-4e49-9a89-125002041b6eCited by top-tier papers6
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen et al.AAAI 2025 · 7 citations
- NegoCollab: A Common Representation Negotiation Approach for Heterogeneous Collaborative PerceptionCongzhang Shao, Quan Yuan, Guiyang Luo, Yue Hu et al.NeurIPS 2025 · 7 citations
- Distilling Cross-Modal Knowledge via Feature DisentanglementJunhong Liu, Yuan Zhang, Tao Huang, Wenchao Xu et al.AAAI 2026 · 2 citations
- LightLoc: Learning Outdoor LiDAR Localization at Light SpeedWen Li, Chen Liu, Shangshu Yu, Dunqiang Liu et al.CVPR 2025
- Multi-Modal Aerial-Ground Cross-View Place Recognition with Neural ODEsSijie Wang, Rui She, Qiyu Kang, Siqi Li et al.CVPR 2025
Builds on9
- Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingMingi Ji, Byeongho Heo, Sungrae ParkAAAI 2021 · 194 citations
- Point-to-Voxel Knowledge Distillation for LiDAR Semantic SegmentationYuenan Hou, Xinge Zhu, Yuexin Ma, Chen Change Loy et al.CVPR 2022 · 185 citations
- Adversarial Robustness in Graph Neural Networks: A Hamiltonian ApproachKai Zhao, Qiyu Kang, Yang Song, Rui She et al.NeurIPS 2023 · 45 citations
- Contextual Similarity Distillation for Asymmetric Image RetrievalHui Wu, Min Wang, Wengang Zhou, Houqiang Li et al.CVPR 2022 · 34 citations
- HypLiLoc: Towards Effective LiDAR Pose Regression with Hyperbolic FusionSijie Wang, Qiyu Kang, Rui She, Wei Wang et al.CVPR 2023
Related papers
- StructVPR: Distill Structural Knowledge with Weighting Samples for Visual Place RecognitionYanqing Shen, Sanping Zhou, Jingwen Fu, Ruotong Wang et al.CVPR 2023
- STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object DetectionSujin Jang, Dae Ung Jo, Sung Ju Hwang, Dongwook Lee et al.NeurIPS 2023 · 19 citations
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou et al.CVPR 2023
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionHaimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen et al.AAAI 2024 · 31 citations
- D²-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable AggregationZheyuan Zhang, Jiwei Zhang, Boyu Zhou, Linzhimeng Duan et al.AAAI 2026 · 2 citations
