Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
Sanghyeob Song, Jaihyun Lew, Hyemi Jang, Sungroh Yoon
Abstract
Estimating the homography between two images is crucial for mid- or high-level vision tasks, such as image stitching and fusion. However, using supervised learning methods is often challenging or costly due to the difficulty of collecting ground-truth data. In response, unsupervised learning approaches have emerged. Most early methods, though, assume that the given image pairs are from the same camera or have minor lighting differences. Consequently, while these methods perform effectively under such conditions, they generally fail when input image pairs come from different domains, referred to as multimodal image pairs. To address these limitations, we propose AltO, an unsupervised learning framework for estimating homography in multimodal image pairs. Our method employs a two-phase alternating optimization framework, similar to Expectation-Maximization (EM), where one phase reduces the geometry gap and the other addresses the modality gap. To handle these gaps, we use Barlow Twins loss for the modality gap and propose an extended version, Geometry Barlow Twins, for the geometry gap. As a result, we demonstrate that our method, AltO, can be trained on multimodal datasets without any ground-truth data. It not only outperforms other unsupervised methods but is also compatible with various architectures of homography estimators. The source code can be found at: https://github.com/songsang7/AltO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f63d6e5-efa7-4c3e-866a-4d344bf1123aCited by top-tier papers3
- Towards Generalized Multimodal Homography EstimationJinkun You, Jiaxin Cheng, Jie Zhang, Yicong ZhouCVPR 2026 · 1 citation
- SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and AudioTong Lu, Zhichun Wang, Yaoyu Zhou, Yiming Guan et al.AAAI 2026
- HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD MapsXuchang Zhong, Xu Cao, Jinke Feng, Hao FangCVPR 2026
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- Iterative Deep Homography EstimationSi-Yuan Cao, Jianxin Hu, Ze-Hua Sheng, Hui-Liang ShenCVPR 2022 · 65 citations
- Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerSi-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu et al.CVPR 2023
Related papers
- SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationJunchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang et al.CVPR 2025
- Unsupervised Homography Estimation with Coplanarity-Aware GANMingbo Hong, Yuhang Lu, Nianjin Ye, Chunyu Lin et al.CVPR 2022 · 62 citations
- Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence ConstraintHai Jiang, Haipeng Li, Yuhang Lu, Songchen Han et al.AAAI 2023 · 19 citations
- SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingYuxin Deng, Botian Wang, Kaining Zhang, Hao Zhang et al.AAAI 2026
- Deep Lucas-Kanade Homography for Multimodal Image AlignmentYiming Zhao, Xinming Huang, Ziming ZhangCVPR 2021
