Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement
Chunlei Zhang, Jiahao Xia, Yun Xiao, Bo Jiang, Jian Zhang
Abstract
Multimodal image registration is a fundamental task and a prerequisite for downstream cross-modal analysis. Despite recent progress in shared feature extraction and multi-scale architectures, two key limitations remain. First, some methods use disentanglement to learn shared features but mainly regularize the shared part, allowing modality-private cues to leak into the shared space. Second, most multi-scale frameworks support only a single transformation type, limiting their applicability when global misalignment and local deformation coexist. To address these issues, we formulate hybrid multimodal registration as jointly learning a stable shared feature space and a unified hybrid transformation. Based on this view, we propose HRNet, a Hybrid Registration Network that couples representation disentanglement with hybrid parameter prediction. A shared backbone with Modality-Specific Batch Normalization (MSBN) extracts multi-scale features, while a Cross-scale Disentanglement and Adaptive Projection (CDAP) module suppresses modality-private cues and projects shared features into a stable subspace for matching. Built on this shared space, a Hybrid Parameter Prediction Module (HPPM) performs non-iterative coarse-to-fine estimation of global rigid parameters and deformation fields, which are fused into a coherent deformation field. Extensive experiments on four multimodal datasets demonstrate state-of-the-art performance on rigid and non-rigid registration tasks. The code is available at the project website 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Iterative Deep Homography EstimationSi-Yuan Cao, Jianxin Hu, Ze-Hua Sheng, Hui-Liang ShenCVPR 2022 · 65 citations
- MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography EstimationHaokai Zhu, Si-Yuan Cao, Jianxin Hu, Sitong Zuo et al.CVPR 2024 · 18 citations
- Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image TranslationMoab Arar, Yiftach Ginger, Dov Danon, Amit H. Bermano et al.CVPR 2020
- Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus TransformerSi-Yuan Cao, Runmin Zhang, Lun Luo, Beinan Yu et al.CVPR 2023
Related papers
- RFNet: Unsupervised Network for Mutually Reinforcing Multi-modal Image Registration and FusionHan Xu, Jiayi Ma, Jiteng Yuan, Zhuliang Le et al.CVPR 2022 · 161 citations
- Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object DetectionMianzhao Wang, Fan Shi, Xu Cheng, Chen Jia et al.CVPR 2026
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu et al.AAAI 2022 · 218 citations
- Learning Multi-Modal Cross-Scale Deformable Transformer Network for Unregistered Hyperspectral Image Super-resolutionWenqian Dong, Yang Xu, Jiahui Qu, Shaoxiong HouAAAI 2024 · 12 citations
- Multi-Modal Object Re-identification via Sparse Mixture-of-ExpertsYingying Feng, Jie Li, Chi Xie, Lei Tan et al.ICML 2025
