Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)
Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, Longbo Huang
摘要
Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network , which is counter-intuitive since multiple signals generally bring more information Wang et al. [2020] . This work provides a theoretical explanation for the emergence of such performance gap in neural networks for the prevalent joint training framework. Based on a simplified data distribution that captures the realistic property of multi-modal data, we prove that for the multi-modal late-fusion network with (smoothed) ReLU activation trained jointly by gradient descent, different modalities will compete with each other. The encoder networks will learn only a subset of modalities. We refer to this phenomenon as modality competition. The losing modalities, which fail to be discovered, are the origins where the sub-optimality of joint training comes from. Experimentally, we illustrate that modality competition matches the intrinsic behavior of late-fusion joint training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper53
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi 等ICML 2023 · 被引用 237 次
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
- In-context Convergence of TransformersYu Huang, Yuan Cheng, Yingbin LiangICML 2024 · 被引用 114 次
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 被引用 86 次
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei 等ICCV 2023 · 被引用 84 次
它引用的顶会 Paper13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
相关 Paper
- Understanding Unimodal Bias in Multimodal Deep Linear NetworksYedi Zhang, Peter E. Latham, Andrew M. SaxeICML 2024 · 被引用 20 次
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 被引用 9 次
- What Makes Training Multi-Modal Classification Networks Hard?Weiyao Wang, Du Tran, Matt FeiszliCVPR 2020
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 被引用 48 次
