Cross-Modal Alignment via Variational Copula Modelling
Feng Wu, Tsai Hor Chan, Fuying Wang, Guosheng Yin, Lequan Yu
Abstract
Various data modalities are common in real-world applications (e.g., electronic health records, medical images and clinical notes in healthcare). It is essential to develop multimodal learning methods to aggregate various information from multiple modalities. The main challenge is how to appropriately align and fuse the representations of different modalities into a joint distribution. Existing methods mainly rely on concatenation or the Kronecker product, oversimplifying the interaction structure between modalities and indicating a need to model more complex interactions. Additionally, the joint distribution of latent representations with higher-order interactions is underexplored. Copula is a powerful statistical structure for modelling the interactions among variables, as it naturally bridges the joint distribution and marginal distributions of multiple variables. We propose a novel copula-driven multimodal learning framework, which focuses on learning the joint distribution of various modalities to capture the complex interactions among them. The key idea is to interpret the copula model as a tool to align the marginal distributions of the modalities efficiently. By assuming a Gaussian mixture distribution for each modality and a copula model on the joint distribution, our model can generate accurate representations for missing modalities. Extensive experiments on public MIMIC datasets demonstrate the superior performance of our model over other competitors. The code is available at https: //github.com/HKU-MedAI/CMCM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cdd24fd-a529-4d54-9bbc-730e3a162876Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling et al.NeurIPS 2023 · 120 citations
- MMTN: Multi-Modal Memory Transformer Network for Image-Report Consistent Medical Report GenerationYiming Cao, Lizhen Cui, Lei Zhang, Fuqiang Yu et al.AAAI 2023 · 56 citations
- Morphological Prototyping for Unsupervised Slide Representation Learning in Computational PathologyAndrew H. Song, Richard J. Chen, Tong Ding, Drew F. K. Williamson et al.CVPR 2024 · 51 citations
Related papers
- Multimodal Learning with Incomplete Modalities by Knowledge DistillationQi Wang, Liang Zhan, Paul M. Thompson, Jiayu ZhouKDD 2020 · 80 citations
- Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisBing Cao, Han Zhang, Nannan Wang, Xinbo Gao et al.AAAI 2020 · 94 citations
- Causal Representation Learning from Multimodal Clinical Records under Non-Random Modality MissingnessZihan Liang, Ziwen Pan, Ruoxuan XiongEMNLP 2025
- Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry AttentionJoy Dhar, Manish Kumar Pandey, Nayyar Zaidi, Chen Chen et al.KDD 2026
- MedAlign: Enhancing Combinatorial Medication Recommendation with Multi-modality AlignmentHang Lv, Zixuan Guo, Zijie Wu, Yanchao Tan et al.ACM MM 2025 · 4 citations
