Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
Sanghyeob Song, Jaihyun Lew, Hyemi Jang, Sungroh Yoon
摘要
Estimating the homography between two images is crucial for mid- or high-level vision tasks, such as image stitching and fusion. However, using supervised learning methods is often challenging or costly due to the difficulty of collecting ground-truth data. In response, unsupervised learning approaches have emerged. Most early methods, though, assume that the given image pairs are from the same camera or have minor lighting differences. Consequently, while these methods perform effectively under such conditions, they generally fail when input image pairs come from different domains, referred to as multimodal image pairs. To address these limitations, we propose AltO, an unsupervised learning framework for estimating homography in multimodal image pairs. Our method employs a two-phase alternating optimization framework, similar to Expectation-Maximization (EM), where one phase reduces the geometry gap and the other addresses the modality gap. To handle these gaps, we use Barlow Twins loss for the modality gap and propose an extended version, Geometry Barlow Twins, for the geometry gap. As a result, we demonstrate that our method, AltO, can be trained on multimodal datasets without any ground-truth data. It not only outperforms other unsupervised methods but is also compatible with various architectures of homography estimators. The source code can be found at: https://github.com/songsang7/AltO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Towards Generalized Multimodal Homography EstimationJinkun You, Jiaxin Cheng, Jie Zhang, Yicong ZhouCVPR 2026 · 被引用 1 次
- SciMKG: A Multimodal Knowledge Graph for Science Education with Text, Image, Video and AudioTong Lu, Zhichun Wang, Yaoyu Zhou, Yiming Guan 等AAAI 2026
- HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD MapsXuchang Zhong, Xu Cao, Jinke Feng, Hao FangCVPR 2026
它引用的顶会 Paper7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- Iterative Deep Homography EstimationSi-Yuan Cao, Jianxin Hu, Ze-Hua Sheng, Hui-Liang ShenCVPR 2022 · 被引用 65 次
- Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerSi-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu 等CVPR 2023
相关 Paper
- SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split OptimizationJunchen Yu, Si-Yuan Cao, Runmin Zhang, Chenghao Zhang 等CVPR 2025
- Unsupervised Homography Estimation with Coplanarity-Aware GANMingbo Hong, Yuhang Lu, Nianjin Ye, Chunyu Lin 等CVPR 2022 · 被引用 62 次
- Semi-supervised Deep Large-Baseline Homography Estimation with Progressive Equivalence ConstraintHai Jiang, Haipeng Li, Yuhang Lu, Songchen Han 等AAAI 2023 · 被引用 19 次
- SGPFeat: Semantic and Geometric Priors for Multi-modal Image MatchingYuxin Deng, Botian Wang, Kaining Zhang, Hao Zhang 等AAAI 2026
- Deep Lucas-Kanade Homography for Multimodal Image AlignmentYiming Zhao, Xinming Huang, Ziming ZhangCVPR 2021
