Understanding and Constructing Latent Modality Structures in Multi-Modal Representation Learning
Qian Jiang, Changyou Chen, Han Zhao, Liqun Chen, Qing Ping, Son Dinh Tran, Yi Xu, Belinda Zeng, Trishul Chilimbi
Abstract
Contrastive loss has been increasingly used in learning representations from multiple modalities. In the limit, the nature of the contrastive loss encourages modalities to exactly match each other in the latent space. Yet it remains an open question how the modality alignment affects the downstream task performance. In this paper, based on an information-theoretic argument, we first prove that exact modality alignment is sub-optimal in general for downstream prediction tasks. Hence we advocate that the key of better performance lies in meaningful latent modality structures instead of perfect modality alignment. To this end, we propose three general approaches to construct latent modality structures. Specifically, we design 1) a deep feature separation loss for intra-modality regularization; 2) a Brownian-bridge loss for inter-modality regularization; and 3) a geometric consistency loss for both intra-and intermodality regularization. Extensive experiments are conducted on two popular multi-modal representation learning frameworks: the CLIP-based two-tower model and the ALBEF-based fusion model. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods, demonstrating the effectiveness and generalizability of our proposed approach on latent modality structure regularization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e939e54-06b3-4122-aecd-dd23ea01b939Cited by top-tier papers32
- SimMMDG: A Simple and Effective Framework for Multi-modal Domain GeneralizationHao Dong, Ismail Nejjar, Han Sun, Eleni N. Chatzi et al.NeurIPS 2023 · 80 citations
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 23 citations
- Cross-Modal Redundancy and the Geometry of Vision-Language EmbeddingsGrégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin M. PicardICLR 2026 · 9 citations
- Learning Disentangled Representation for Multi-Modal Time-Series Sensing SignalsRuichu Cai, Zhifan Jiang, Kaitao Zheng, Zijian Li et al.WWW 2025 · 8 citations
- Mitigating Noisy Correspondence by Geometrical Structure Consistency LearningZihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao et al.CVPR 2024 · 5 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Understanding Transferable Representation Learning and Zero-shot Transfer in CLIPZixiang Chen, Yihe Deng, Yuanzhi Li, Quanquan GuICLR 2024 · 21 citations
- Vision-Language Pre-Training with Triple Contrastive LearningJinyu Yang, Jiali Duan, Son Tran, Yi Xu et al.CVPR 2022 · 266 citations
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini et al.ICLR 2025
- Aligning Multimodal Representations through an Information BottleneckAntonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer et al.ICML 2025
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 1 citation
