A Theoretical Proof of Dynamic Multimodal Fusion Exacerbates Modality Greedy
Xiaorui Ding, Huan Ma, Changqing Zhang
Abstract
In recent years, numerous studies have proposed uncertainty-guided multimodal learning to adapt to dynamic relationships between different modalities. Dynamic multimodal fusion enables more dominant modalities to receive greater weight during the fusion process, thereby preventing the influence of spurious features from less reliable modalities on decision-making. However, there is No free lunch. We observe that the introduction of dynamic fusion during training exacerbates the model's tendency toward Greedy (a phenomenon known to induce decision shortcuts in multimodal learning). This results in a model that does not fully take advantage of the lower quality modalities. In this paper, we provide a theoretical analysis showing that dynamic fusion intensifies Greedy, and we present experimental results that support this observation. In summary, this paper explains the Greedy risk in dynamic multimodal learning from both theoretical and experimental perspectives, serving as a cautionary reminder for researchers when employing dynamic multimodal learning. Our code is available at https://github.com/d-xr/GreedyDynFusion.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 98630caf-4da0-4dea-a571-2bb309f3af8fRelated papers
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- CL-DMDF: Dynamic Multimodal Data Fusion Model Based on Contrastive LearningDong Li, Lingling Zhang, Binghao Han, Linlin Ding et al.AAAI 2026
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 1 citation
- Predictive Dynamic FusionBing Cao, Yinan Xia, Yi Ding, Changqing Zhang et al.ICML 2024 · 31 citations
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 9 citations
