Aligning Multimodal Representations through an Information Bottleneck
Antonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer, Alfonso Ortega
摘要
Contrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-specific information in the representation space. Although some of the most widely used contrastive losses maximize the mutual information between representations of both modalities, they are not designed to remove the modalityspecific information. We give a theoretical description of this problem through the lens of the Information Bottleneck Principle. We also empirically analyze how different hyperparameters affect the emergence of this phenomenon in a controlled experimental setup. Finally, we propose a regularization term in the loss function that is derived by means of a variational approximation and aims to increase the representational alignment. We analyze in a set of controlled experiments and real-world applications the advantages of including this regularization term.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ECHO: Toward Contextual Seq2Seq Paradigms in Large EEG ModelsChenyu Liu, Yuqiu Deng, Tianyu Liu, Jinan Zhou 等ICLR 2026 · 被引用 12 次
- Distributional Vision-Language Alignment by Cauchy-Schwarz DivergenceWenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu 等ICLR 2026 · 被引用 9 次
- A Comprehensive Information-Decomposition Analysis of Large Vision-Language ModelsLixin Xiu, Xufang Luo, Hideki NakayamaICLR 2026 · 被引用 4 次
- Bridging Functional and Representational Similarity via Usable InformationAntonio Almudévar, Alfonso OrtegaICML 2026 · 被引用 1 次
- Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-ExpertsHahyeon Choi, NOJUN KWAKICML 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva 等ICML 2026
- Learning Optimal Multimodal Information Bottleneck RepresentationsQilong Wu, Yiyang Shao, Jun Wang, Xiaobo SunICML 2025
- IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity AlignmentTaoyu Su, Jiawei Sheng, Shicheng Wang, Xinghua Zhang 等ACM MM 2024 · 被引用 7 次
- To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal PerformanceWanlong Fang, Tianle Zhang, Alvin ChanAAAI 2026
- Understanding and Constructing Latent Modality Structures in Multi-Modal Representation LearningQian Jiang, Changyou Chen, Han Zhao, Liqun Chen 等CVPR 2023
