A Theory of Multimodal Learning
Zhou Lu
摘要
Human perception of the empirical world involves recognizing the diverse appearances, or 'modalities', of underlying objects. Despite the longstanding consideration of this perspective in philosophy and cognitive science, the study of multimodality remains relatively under-explored within the field of machine learning. Nevertheless, current studies of multimodal machine learning are limited to empirical practices, lacking theoretical foundations beyond heuristic arguments. An intriguing finding from the practice of multimodal learning is that a model trained on multiple modalities can outperform a finely-tuned unimodal model, even on unimodal tasks. This paper provides a theoretical framework that explains this phenomenon, by studying generalization properties of multimodal learning algorithms. We demonstrate that multimodal learning allows for a superior generalization bound compared to unimodal learning, up to a factor of , where represents the sample size. Such advantage occurs when both connection and heterogeneity exist between the modalities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual GenerationGwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou 等NeurIPS 2024 · 被引用 23 次
- Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variablesYu Gui, Cong Ma, Zongming MaNeurIPS 2025 · 被引用 9 次
- Disentangled Cross-Modal Representation Learning with Enhanced Mutual SupervisionLu Gao, Wenlan Chen, Daoyuan Wang, Fei Guo 等NeurIPS 2025 · 被引用 5 次
- Calibrated Multimodal Representation Learning with Missing ModalitiesXiaohao Liu, Xiaobo Xia, Jiaheng Wei, Shuo Yang 等ICML 2026 · 被引用 5 次
- CMoB: Modality Valuation via Causal Effect for Balanced Multimodal LearningJun Wang, Fuyuan Cao, Zhixin Xue, Xingwang Zhao 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper5
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 被引用 263 次
- On the Provable Advantage of Unsupervised PretrainingJiawei Ge, Shange Tang, Jianqing Fan, Chi JinICLR 2024 · 被引用 23 次
相关 Paper
- Understanding Multimodal Learning: A Loss Landscape Smoothness PerspectiveJae-Jun Lee, Sung Whan YoonICML 2026
- Understanding the Emergence of Multimodal Representation AlignmentMegan Tjandrasuwita, Chanakya Ekbote, Liu Ziyin, Paul Pu LiangICML 2025
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot TasksXizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu 等CVPR 2022
- Towards Out-of-Modal Generalization without Instance-level Modal CorrespondenceZhuo Huang, Gang Niu, Bo Han, Masashi Sugiyama 等ICLR 2025
