What to align in multimodal contrastive learning?
Benoit Dufumier, Javiera Castillo Navarro, Devis Tuia, Jean-Philippe Thiran
Abstract
Humans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior. Alignment through contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive Multimodal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross-or intra-modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, we show that CoMM learns complex multimodal interactions and achieves state-of-the-art results on seven multimodal tasks. Code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa31c758-8320-4b5d-a903-2f33e677e2bdCited by top-tier papers20
- COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Wei Wang, Xiping Hu et al.SIGIR 2025 · 20 citations
- InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsLiangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng et al.NeurIPS 2025 · 10 citations
- Partial Information Decomposition via Normalizing Flows in Latent Gaussian DistributionsWenyuan Zhao, Adithya Balachandran, Chao Tian, Paul Pu LiangNeurIPS 2025 · 5 citations
- THE MORE, THE MERRIER: CONTRASTIVE FUSION FOR HIGHER-ORDER MULTIMODAL ALIGNMENTStefanos Koutoupis, Michaela Areti Zervou, Konstantinos Kontras, Maarten De Vos et al.CVPR 2026 · 5 citations
- Closing the Modality Gap Aligns Group-Wise SemanticsEleonora Grassucci, Giordano Cicchetti, Emanuele Frasca, Aurelio Uncini et al.ICLR 2026 · 5 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Contrastive Multimodal Fusion with TupleInfoNCEYunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong et al.ICCV 2021 · 84 citations
- To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal PerformanceWanlong Fang, Tianle Zhang, Alvin ChanAAAI 2026
- Information-Theoretic Decomposition for Multimodal Interaction LearningZequn Yang, Yake Wei, Haotian Ni, Zhihao Xu et al.CVPR 2026 · 1 citation
- Mutual Contrastive Learning for Visual Representation LearningChuanguang Yang, Zhulin An, Linhang Cai, Yongjun XuAAAI 2022 · 95 citations
- Multimodal Classification via Total Correlation MaximizationFeng Yu, Xiangyu Wu, Yang Yang, Jianfeng LuICLR 2026 · 4 citations
