To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
Wanlong Fang, Tianle Zhang, Alvin Chan
Abstract
Multimodal learning often relies on aligning representations across modalities to enable effective information integration—an approach traditionally assumed to be universally beneficial. However, prior research has primarily taken an observational approach, examining naturally occurring alignment in multimodal data and exploring its correlation with model performance, without systematically studying the direct effects of explicitly enforced alignment between representations of different modalities. In this work, we investigate how explicit alignment influences both model performance and representation alignment under different modality-specific information structures. Specifically, we introduce a controllable contrastive learning module that enables precise manipulation of alignment strength during training, allowing us to explore when explicit alignment improves or hinders performance. Our results on synthetic and real datasets under different data characteristics show that the impact of explicit alignment on the performance of unimodal models is related to the characteristics of the data: the optimal level of alignment depends on the amount of redundancy between the different modalities. We can find an optimal alignment strength that balances modality-specific signals and shared redundancy in the mixed information distributions. This work can help practitioners on when and how to enforce alignment for optimal unimodal encoder performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3ae1732-97b7-4bf7-91c6-68c97aaa5fddCited by top-tier papers7
- Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language NavigationXiang Fang, Wanlong Fang, Changshuo WangNeurIPS 2025 · 24 citations
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information DecompositionWanlong Fang, Tianle Zhang, Wen Tao, Alvin ChanICML 2026 · 17 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu et al.ACM MM 2024 · 8 citations
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM SecurityXiang Fang, Wanlong FangAAAI 2026 · 4 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video RepresentationsMohammadreza Zolfaghari, Yi Zhu, Peter V. Gehler, Thomas BroxICCV 2021 · 160 citations
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling et al.NeurIPS 2023 · 120 citations
Related papers
- Understanding the Emergence of Multimodal Representation AlignmentMegan Tjandrasuwita, Chanakya Ekbote, Liu Ziyin, Paul Pu LiangICML 2025
- Aligning Multimodal Representations through an Information BottleneckAntonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer et al.ICML 2025
- On the Value of Cross-Modal Misalignment in Multimodal Representation LearningYichao Cai, Yuhang Liu, Erdun Gao, Tianjiao Jiang et al.NeurIPS 2025 · 11 citations
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva et al.ICML 2026
- What to align in multimodal contrastive learning?Benoit Dufumier, Javiera Castillo Navarro, Devis Tuia, Jean-Philippe ThiranICLR 2025 · 3 citations
