What Makes Multi-Modal Learning Better than Single (Provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, Longbo Huang
Abstract
The world provides us with data of multiple modalities. Intuitively, models fusing data from different modalities outperform their uni-modal counterparts, since more information is aggregated. Recently, joining the success of deep learning, there is an influential line of work on deep multi-modal learning, which has remarkable empirical results on various applications. However, theoretical justifications in this field are notably lacking. Can multi-modal learning provably perform better than uni-modal? In this paper, we answer this question under a most popular multi-modal fusion framework, which firstly encodes features from different modalities into a common latent space and seamlessly maps the latent representations into the task space. We prove that learning with multiple modalities achieves a smaller population risk than only using its subset of modalities. The main intuition is that the former has a more accurate estimate of the latent space representation. To the best of our knowledge, this is the first theoretical treatment to capture important qualitative phenomena observed in real multi-modal applications from the generalization perspective. Combining with experiment results, we show that multi-modal learning does possess an appealing formal guarantee.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49673eb2-465a-4d9d-bd0b-22960c0f4e2eCited by top-tier papers70
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang et al.CVPR 2022 · 149 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- Factorized Contrastive Learning: Going Beyond Multi-view RedundancyPaul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou et al.NeurIPS 2023 · 137 citations
- Divert More Attention to Vision-Language TrackingMingzhe Guo, Zhipeng Zhang, Heng Fan, Liping JingNeurIPS 2022 · 122 citations
Builds on9
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
- Self-supervised Learning from a Multi-view PerspectiveYao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, Louis-Philippe MorencyICLR 2021 · 232 citations
- Provable Meta-Learning of Linear RepresentationsNilesh Tripuraneni, Chi Jin, Michael I. JordanICML 2021 · 218 citations
Related papers
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 48 citations
- Efficient Multi-Modal Fusion with Diversity AnalysisShuhui Qu, Yan Kang, Janghwan LeeACM MM 2021 · 4 citations
- Multimodal Learning with Incomplete Modalities by Knowledge DistillationQi Wang, Liang Zhan, Paul M. Thompson, Jiayu ZhouKDD 2020 · 80 citations
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu et al.ICML 2023 · 79 citations
- Safe Multi-View Deep ClassificationWei Liu, Yufei Chen, Xiaodong Yue, Changqing Zhang et al.AAAI 2023 · 27 citations
