Can One Modality Model Synergize Training of Other Modality Models?
Jae-Jun Lee, Sung Whan Yoon
摘要
Learning with multiple modalities has recently demonstrated significant gains in many domains by maximizing the shared information across modalities. However, the current approaches strongly rely on high-quality paired datasets, which allow co-training from the paired labels from different modalities. In this context, we raise a pivotal question: Can a model with one modality synergize the training of other models with the different modalities, even without the paired multimodal supervision? Our answer is 'Yes'. As a figurative description, we argue that a writer, i.e., a language model, can promote the training of a painter, i.e., a visual model, even without the paired ground truth of text and image. We theoretically show that a superior representation can be achieved by the synergy between two different modalities, without paired supervision. As proofs of concept, we broadly confirm the considerable gains from the synergy across visual, language, and audio models. From a theoretical viewpoint, we first establish a mathematical foundation of the synergy between two different modality models, where each one is trained with its own modality. From a practical viewpoint, our work aims to broaden the scope of multimodal learning to encompass the synergistic usage of single-modality models, relieving a strong limitation of paired supervision. The code is available at https://github.com/johnjaejunlee95/synergistic-multimodal .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal ModelsSharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka 等ICLR 2026
- Understanding Multimodal Learning: A Loss Landscape Smoothness PerspectiveJae-Jun Lee, Sung Whan YoonICML 2026
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
相关 Paper
- Autoregressive Pre-Training on Pixels and TextsYekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang 等EMNLP 2024 · 被引用 1 次
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic 等NeurIPS 2020 · 被引用 423 次
- Fusing Pre-Trained Language Models with Multimodal Prompts through Reinforcement LearningYoungjae Yu, Jiwan Chung, Heeseung Yun, Jack Hessel 等CVPR 2023
- Non-Linguistic Supervision for Contrastive Learning of Sentence EmbeddingsYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2022 · 被引用 20 次
- Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech RecognitionXichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou 等ACL 2022 · 被引用 43 次
