Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)
Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, Longbo Huang
Abstract
Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network , which is counter-intuitive since multiple signals generally bring more information Wang et al. [2020] . This work provides a theoretical explanation for the emergence of such performance gap in neural networks for the prevalent joint training framework. Based on a simplified data distribution that captures the realistic property of multi-modal data, we prove that for the multi-modal late-fusion network with (smoothed) ReLU activation trained jointly by gradient descent, different modalities will compete with each other. The encoder networks will learn only a subset of modalities. We refer to this phenomenon as modality competition. The losing modalities, which fail to be discovered, are the origins where the sub-optimality of joint training comes from. Experimentally, we illustrate that modality competition matches the intrinsic behavior of late-fusion joint training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85e1c3be-40fb-4bcb-9a10-c00963ee5e9dCited by top-tier papers53
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi et al.ICML 2023 · 237 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- In-context Convergence of TransformersYu Huang, Yuan Cheng, Yingbin LiangICML 2024 · 114 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei et al.ICCV 2023 · 84 citations
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
Related papers
- Understanding Unimodal Bias in Multimodal Deep Linear NetworksYedi Zhang, Peter E. Latham, Andrew M. SaxeICML 2024 · 20 citations
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu et al.ICML 2023 · 79 citations
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 9 citations
- What Makes Training Multi-Modal Classification Networks Hard?Weiyao Wang, Du Tran, Matt FeiszliCVPR 2020
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 48 citations
