Disentangling for Transfer: Boosting Limited Modalities via Information-Theoretic Regularization and Cross-Modal Reconstruction
Zhiyun Zhang, Yan-Jie Zhou, Yujian Hu, Xiyao Ma, Zhouhang Yuan, Zirui Wang, Hongkun Zhang, Minfeng Xu
Abstract
Missing critical modalities in medical imaging poses significant challenges for AI-driven diagnostic systems, particularly in scenarios where limited modalities must suffice for downstream tasks. Existing approaches often fail to fully leverage privileged features available only at training or address the information gap between privileged and limited modalities, resulting in suboptimal performance. To address this, we propose a unified, dual-stage Disentanglement-AligNmenT framEwork (DANTE), which uses InformationTheoretic Regularization and Cross-Modal Reconstruction to decompose full-modality information into alignable and privileged-exclusive components. In the first stage, a self-supervised pre-training strategy based on cross-modal reconstruction acts as a proxy task to implicitly incentivize disentangled representations. In the second stage, we present an information-theoretic regularization to explicitly maximize the transfer of privileged knowledge through two novel modules: (1) a Mutual Alignment Module that employs multilevel bidirectional alignment between limited-modality features and alignable features, enhancing cross-modal representation consistency; (2) a Privileged Compaction Module that restricts the privileged-exclusive information flow, promoting the integration of task-relevant content into alignable representations. Experimental results on three challenging medical datasets demonstrate that DANTE achieves state-of-the-art performance, demonstrating its effectiveness in leveraging privileged guidance under modality scarcity, and exhibits broad applicability across diverse medical imaging scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 208 citations
- Modal-aware Visual Prompting for Incomplete Multi-modal Brain Tumor SegmentationYansheng Qiu, Ziyuan Zhao, Hongdou Yao, Delin Chen et al.ACM MM 2023 · 25 citations
- C2KD: Bridging the Modality Gap for Cross-Modal Knowledge DistillationFushuo Huo, Wenchao Xu, Jingcai Guo, Haozhao Wang et al.CVPR 2024
- CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-TrainingYuxin Guo, Siyang Sun, Shuailei Ma, Kecheng Zheng et al.CVPR 2024
- MMANet: Margin-Aware Distillation and Modality-Aware Regularization for Incomplete Multimodal LearningShicai Wei, Chunbo Luo, Yang LuoCVPR 2023
Related papers
- Multi-modal Vision Pre-training for Medical Image AnalysisShaohao Rui, Lingzhi Chen, Zhenyu Tang, Lilong Wang et al.CVPR 2025
- Incomplete Modality Disentangled Representation for Ophthalmic Disease Grading and DiagnosisChengzhi Liu, Zile Huang, Zhe Chen, Feilong Tang et al.AAAI 2025 · 10 citations
- Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality AdaptationYunpeng Zhao, Cheng Chen, Qing You Pang, Yibing Fu et al.AAAI 2026
- SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentationkaiwen Huang, Yi Zhou, Yizhe Zhang, Jingxiong Li et al.CVPR 2026
- KAMP: Knowledge-Anchored Multimodal Pretraining Framework for Medical Image RepresentationFeiyu Huang, Jia Li, Zhao Chen, Yang Wu et al.CVPR 2026
