Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models
Sharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka, Phillip Isola
Abstract
Traditional multimodal learners find unified representations for tasks like visual question answering, but rely heavily on large paired datasets. However, an overlooked yet potentially powerful question is: can one leverage auxiliary multimodal data to directly enhance representation learning in a modality? We introduce : npaired ultimodal earner, a modality-agnostic training paradigm in which a single model alternately processes inputs from different modalities while sharing parameters across them. This design exploits the assumption that different modalities are projections of a shared underlying reality, allowing the model to benefit from cross-modal structure without requiring explicit pairs. Theoretically, under linear data-generating assumptions, we show that unpaired auxiliary data can yield representations strictly more informative about the world than unimodal training. Empirically, we show that incorporating unpaired data that share underlying semantic information from auxiliary modalities—such as text, audio, or images—consistently improves downstream performance across diverse unimodal targets such as image and audio. Our project page: https://unpaired-multimodal.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 109896b7-35d5-416b-99ad-d06f0762f87bCited by top-tier papers3
- Calibrated Multimodal Representation Learning with Missing ModalitiesXiaohao Liu, Xiaobo Xia, Jiaheng Wei, Shuo Yang et al.ICML 2026 · 5 citations
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the EdgeYu Zhang, Xi Zhang, Hualin zhou, Xinyuan Chen et al.ICML 2026
- Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent ActionsYongqi Li, Hao Lang, Tieyun Qian, Yongbin LiACL 2026
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Can One Modality Model Synergize Training of Other Modality Models?Jae-Jun Lee, Sung Whan YoonICLR 2025
- Multimodal Pathway: Improve Transformers with Irrelevant Data from Other ModalitiesYiyuan Zhang, Xiaohan Ding, Kaixiong Gong, Yixiao Ge et al.CVPR 2024 · 6 citations
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- The "Law" of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired ModalitiesYongwei Che, Benjamin EysenbachICLR 2025
