Understanding the Emergence of Multimodal Representation Alignment
Megan Tjandrasuwita, Chanakya Ekbote, Liu Ziyin, Paul Pu Liang
Abstract
Multimodal representation learning is fundamentally about transforming incomparable modalities into comparable representations. While prior research primarily focused on explicitly aligning these representations through targeted learning objectives and model architectures, a recent line of work has found that independently trained unimodal models of increasing scale and performance can become implicitly aligned with each other. These findings raise fundamental questions regarding the emergence of aligned representations in multimodal learning. Specifically: (1) when and why does alignment emerge implicitly? and ( 2 ) is alignment a reliable indicator of performance? Through a comprehensive empirical investigation, we demonstrate that both the emergence of alignment and its relationship with task performance depend on several critical data characteristics. These include, but are not necessarily limited to, the degree of similarity between the modalities and the balance between redundant and unique information they provide for the task. Our findings suggest that alignment may not be universally beneficial; rather, its impact on performance varies depending on the dataset and task. These insights can help practitioners determine whether increasing alignment between modalities is advantageous or, in some cases, detrimental to achieving optimal performance. Code is released at: https://github. com/MeganTj/multimodal_alignment .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70c41678-7d07-41f9-8ea7-05fbbf50520eCited by top-tier papers9
- Revisiting the Platonic Representation Hypothesis: An Aristotelian ViewFabian Gröger, Shuo Wen, Maria BrbicICML 2026 · 27 citations
- Neural Thermodynamics: Entropic Forces in Deep and Universal Representation LearningLiu Ziyin, Yizhou Xu, Isaac L. ChuangNeurIPS 2025 · 11 citations
- Dynamic Reflections: Probing Video Representations with Text AlignmentMaks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu et al.ICLR 2026 · 5 citations
- Calibrated Multimodal Representation Learning with Missing ModalitiesXiaohao Liu, Xiaobo Xia, Jiaheng Wei, Shuo Yang et al.ICML 2026 · 5 citations
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalDivyam Madaan, Varshan Muhunthan, Kyunghyun Cho, Sumit ChopraICLR 2026 · 3 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal PerformanceWanlong Fang, Tianle Zhang, Alvin ChanAAAI 2026
- Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal ModelsSharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka et al.ICLR 2026
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva et al.ICML 2026
- DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation LearningChengxuan Qian, Shuo Xing, Li Li, Yue Zhao et al.ICLR 2026 · 42 citations
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 48 citations
