Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention
Joy Dhar, Manish Kumar Pandey, Nayyar Zaidi, Chen Chen, Maryam Haghighat, Ferdous Sohel, Puneet Goyal
Abstract
Multimodal fusion learning (MFL) (a framework to jointly learn from heterogeneous data sources) has shown great potential in various fields such as Medicine, Science, and Engineering. It is extremely desirable in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively, which in turn limits performance improvements. Second, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI applications. Finally, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. To address these challenges, we propose a novel MFL framework – Cascaded Unified Representation Learning for Efficient Fusion Network (CURE) – a lightweight and scalable framework that progressively integrates various modalities through a novel efficient Hybrid Geometry Aware Fusion layer (HyFuse), where each HyFuse layer is sequentially learned for each modality, making the framework adaptable and generalizable. Within HyFuse, an efficient residual convolution module captures rich multi-scale features to ensure cost-effective learning, while a hybrid-space aware attention mixer learns coarse-to-fine structural cues to better preserve cross-modal relationships. Complementary learnable late-fusion and shared-information refinement modules are then employed to learn robust, modality-order-invariant shared features, which in turn yields consistent performance improvements. Extensive evaluations on 16 public datasets show that CURE outperforms leading multimodal fusion methods (e.g., DRIFA-Net and HEALNet), boosting performance by up to ≈ 3.97% and lowering computational costs by up to ≈ 87.8%, ensuring more effective and reliable predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-efficient Medical Image RecognitionShih-Cheng Huang, Liyue Shen, Matthew P. Lungren, Serena YeungICCV 2021 · 516 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- EMCAD: Efficient Multi-Scale Convolutional Attention Decoding for Medical Image SegmentationMd Mostafijur Rahman, Mustafa Munir, Radu MarculescuCVPR 2024 · 352 citations
Related papers
- HEALNet: Multimodal Fusion for Heterogeneous Biomedical DataKonstantin Hemker, Nikola Simidjievski, Mateja JamnikNeurIPS 2024 · 80 citations
- Effective and Robust Multimodal Medical Image AnalysisJoy Dhar, Nayyar Zaidi, Maryam HaghighatKDD 2026 · 1 citation
- Multi-modal Medical Diagnosis via Large-small Model CollaborationWanyi Chen, Zihua Zhao, Jiangchao Yao, Ya Zhang et al.CVPR 2025
- Cross-Modal Alignment via Variational Copula ModellingFeng Wu, Tsai Hor Chan, Fuying Wang, Guosheng Yin et al.ICML 2025
- DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal InconsistencyWenfang Yao, Kejing Yin, William K. Cheung, Jia Liu et al.AAAI 2024 · 80 citations
