What Makes Multi-Modal Learning Better than Single (Provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, Longbo Huang
摘要
The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is aggregated. Recently, joining the success of deep learning, there is an influential line of work on deep multi-modal learning, which has remarkable empirical results on various applications. However, theoretical justifications in this field are notably lacking. Can multi-modal learning provably perform better than uni-modal? In this paper, we answer this question under a most popular multi-modal fusion framework, which firstly encodes features from different modalities into a common latent space and seamlessly maps the latent representations into the task space. We prove that learning with multiple modalities achieves a smaller population risk than only using its subset of modalities. The main intuition is that the former has a more accurate estimate of the latent space representation. To the best of our knowledge, this is the first theoretical treatment to capture important qualitative phenomena observed in real multi-modal applications from the generalization perspective. Combining with experiment results, we show that multi-modal learning does possess an appealing formal guarantee.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper70
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang 等CVPR 2022 · 被引用 149 次
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou 等NeurIPS 2023 · 被引用 137 次
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 被引用 122 次
它引用的顶会 Paper9
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu 等NeurIPS 2020 · 被引用 321 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
- Self-supervised Learning from a Multi-view PerspectiveYao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, Louis-Philippe MorencyICLR 2021 · 被引用 232 次
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 被引用 218 次
相关 Paper
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 被引用 48 次
- Efficient Multi-Modal Fusion with Diversity AnalysisShuhui Qu, Yan Kang, Janghwan LeeACM MM 2021 · 被引用 4 次
- Multimodal Learning with Incomplete Modalities by Knowledge DistillationQi Wang, Liang Zhan, Paul M. Thompson, Jiayu ZhouKDD 2020 · 被引用 80 次
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- Safe Multi-View Deep ClassificationWei Liu, Yufei Chen, Xiaodong Yue, Changqing Zhang 等AAAI 2023 · 被引用 27 次
