Supporting Multimodal Intermediate Fusion with Informatic Constraint and Distribution Coherence
Yi Li, Fei Song, Changwen Zheng, Jiangmeng Li
摘要
Based on the prevalent intermediate fusion (IF) and late fusion (LF) frameworks, multimodal representation learning (MML) demonstrates its superiority over unimodal representation learning. To investigate the intrinsic factors underlying the empirical success of MML, research grounded in theoretical justifications from the perspective of generalization error has emerged. However, these provable MML studies derive the theoretical findings based on LF, while theoretical exploration based on IF remains scarce. This naturally gives rise to a question: Can we design a comprehensive MML approach supported by the sufficient theoretical analysis across fusion types? To this end, we revisit the IF and LF paradigms from a fine-grained dimensional perspective. The derived theoretical evidence sufficiently establishes the superiority of IF over LF under a specific constraint. Based on a general -Lipschitz continuity assumption, we derive the generalization error upper bound of the IF-based methods, indicating that eliminating the distribution incoherence can improve the generalizability of IF-based MML methods. Building upon these theoretical insights, we establish a novel IF-based MML method, which introduces the informatic constraint and performs distribution cohering. Extensive experimental results on multiple widely adopted datasets verify the effectiveness of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman 等ICLR 2020 · 被引用 330 次
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu 等NeurIPS 2020 · 被引用 321 次
相关 Paper
- Understanding Multimodal Learning: A Loss Landscape Smoothness PerspectiveJae-Jun Lee, Sung Whan YoonICML 2026
- Scaling Language-centric Omnimodal Representation LearningChenghao Xiao, Hou Pong Chan, Hao Zhang, Weiwen Xu 等NeurIPS 2025 · 被引用 25 次
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian 等CVPR 2026 · 被引用 9 次
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- InfoBridge: Balanced Multimodal Integration through Conditional Dependency ModelingChenxin Li, Yifan Liu, Panwang Pan, Hengyu Liu 等ICCV 2025 · 被引用 1 次
