SGPFeat: Semantic and Geometric Priors for Multi-modal Image Matching
Yuxin Deng, Botian Wang, Kaining Zhang, Hao Zhang, Jiayi Ma
Abstract
Multi-modal image matching is a fundamental task in multi-view and multi-modal image processing. Its key challenge lies in extracting features that remain consistent despite drastic appearance variations across modalities. However, the learning of the feature is hindered by the scarcity and the inaccurate alignment of existing multi-modal datasets. To address this, we propose a knowledge distillation framework termed SGPFeat that transfers rich prior knowledge from large-scale unimodal tasks to enhance multi-modal representation learning. Specifically, semantic priors from a vision foundation model guide the feature extractor to identify shared semantic structures across modalities, enabling better generalization under large appearance gaps. In parallel, geometric priors derived from accurately aligned visible-light datasets improve detection precision on noisy aligned multi-modal pairs. Furthermore, we introduce a Heterogeneous Feature Aggregation (HFA) module to facilitate effective distillation and feature representation. Extensive experiments demonstrate that semantic and geometric priors bring significant improvement for our SGPFeat across diverse multi-modal image matching benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ca7b841-02de-4c8d-ab4e-9ac52674299fBuilds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- XFeat: Accelerated Features for Lightweight Image MatchingGuilherme A. Potje, Felipe Cadar, André Araújo, Renato Martins et al.CVPR 2024 · 128 citations
Related papers
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionHaimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen et al.AAAI 2024 · 31 citations
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan et al.CVPR 2021
- Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and AlgorithmT. K Tran, Duc Chu Anh, Quang Hung Pham, Phi Le Nguyen et al.ICML 2026
- Semantic-Guided Feature Distillation for Multimodal RecommendationFan Liu, Huilin Chen, Zhiyong Cheng, Liqiang Nie et al.ACM MM 2023 · 24 citations
- On Modality Weighting and Specificity for Multi-Modal Entity AlignmentYu Xing, Qizhuo Xie, Yunhui Liu, Qing Gu et al.AAAI 2026
