A Theoretical Proof of Dynamic Multimodal Fusion Exacerbates Modality Greedy
Xiaorui Ding, Huan Ma, Changqing Zhang
摘要
In recent years, numerous studies have proposed uncertainty-guided multimodal learning to adapt to dynamic relationships between different modalities. Dynamic multimodal fusion enables more dominant modalities to receive greater weight during the fusion process, thereby preventing the influence of spurious features from less reliable modalities on decision-making. However, there is No free lunch. We observe that the introduction of dynamic fusion during training exacerbates the model's tendency toward Greedy (a phenomenon known to induce decision shortcuts in multimodal learning). This results in a model that does not fully take advantage of the lower quality modalities. In this paper, we provide a theoretical analysis showing that dynamic fusion intensifies Greedy, and we present experimental results that support this observation. In summary, this paper explains the Greedy risk in dynamic multimodal learning from both theoretical and experimental perspectives, serving as a cautionary reminder for researchers when employing dynamic multimodal learning. Our code is available at https://github.com/d-xr/GreedyDynFusion.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding 等AAAI 2026
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 被引用 1 次
- Predictive Dynamic FusionBing Cao, Yinan Xia, Yi Ding, Changqing Zhang 等ICML 2024 · 被引用 31 次
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 被引用 9 次
