Sparse-to-dense Multimodal Image Registration via Multi-Task Learning
Kaining Zhang, Jiayi Ma
Abstract
Aligning image pairs captured by different sensors or those undergoing significant appearance changes is crucial for various computer vision and robotics applications. Existing approaches cope with this problem via either Sparse feature Matching (SM) or Dense direct Alignment (DA) paradigms. Sparse methods are efficient but lack accuracy in textureless scenes, while dense ones are more accurate in all scenes but demand for good initialization. In this paper, we propose SDME, a Sparse-to-Dense Multimodal feature Extractor based on a novel multi-task network that simultaneously predicts SM and DA features for robust multimodal image registration. We propose the sparse-to-dense registration paradigm: we first perform initial registration via SM and then refine the result via DA. By using the welldesigned SDME, the sparse-to-dense approach combines the merits from both SM and DA. Extensive experiments on MSCOCO, GoogleEarth, VIS-NIR and VIS-IR-drone datasets demonstrate that our method achieves remarkable performance on multimodal cases. Furthermore, our approach exhibits robust generalization capabilities, enabling the fine-tuning of models initially trained on single-modal datasets for use with smaller multimodal datasets. Our code is available at https: //github.com/KN-Zhang/SDME .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0c9de8d-8b49-4b00-af2a-cd07b12dbe8aCited by top-tier papers2
- MINIMA: Modality Invariant Image MatchingJiangwei Ren, Xingyu Jiang, Zizhuo Li, Dingkang Liang et al.CVPR 2025
- Adapting Dense Matching for Homography Estimation with Grid-based AccelerationKaining Zhang, Yuxin Deng, Jiayi Ma, Paolo FavaroCVPR 2025
Builds on11
- Which Tasks Should Be Learned Together in Multi-task Learning?Trevor Standley, Amir Zamir, Dawn Chen, Leonidas J. Guibas et al.ICML 2020 · 651 citations
- Learning to Branch for Multi-Task LearningPengsheng Guo, Chen-Yu Lee, Daniel UlbrichtICML 2020 · 208 citations
- Iterative Deep Homography EstimationSi-Yuan Cao, Jianxin Hu, Ze-Hua Sheng, Hui-Liang ShenCVPR 2022 · 65 citations
- Learning Super-Features for Image RetrievalPhilippe Weinzaepfel, Thomas Lucas, Diane Larlus, Yannis KalantidisICLR 2022 · 56 citations
- Cross-Modal Contrastive Learning for Domain Adaptation in 3D Semantic SegmentationBowei Xing, Xianghua Ying, Ruibin Wang, Jinfa Yang et al.AAAI 2023 · 23 citations
Related papers
- RGB-Multispectral Matching: Dataset, Learning Methodology, EvaluationFabio Tosi, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti et al.CVPR 2022 · 5 citations
- SM3Det: A Unified Model for Multi-Modal Remote Sensing Object DetectionYuxuan Li, Xiang Li, Yunheng Li, Yicheng Zhang et al.AAAI 2026 · 25 citations
- Differentiable Registration of Images and LiDAR Point Clouds with VoxelPoint-to-Pixel MatchingJunsheng Zhou, Baorui Ma, Wenyuan Zhang, Yi Fang et al.NeurIPS 2023 · 62 citations
- CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image RegistrationXuecong Liu, Mengzhu Ding, Zixuan Sun, Zhang Li et al.CVPR 2026 · 4 citations
- Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image TranslationMoab Arar, Yiftach Ginger, Dov Danon, Amit H. Bermano et al.CVPR 2020
