Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks
Nan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. Geras
Abstract
We hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities. Such behavior is counter-intuitive and hurts the models' generalization, as we observe empirically. To estimate the model's dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality. We refer to this gain as the conditional utilization rate. In the experiments, we consistently observe an imbalance in conditional utilization rates between modalities, across multiple tasks and architectures. Since conditional utilization rate cannot be computed efficiently during training, we introduce a proxy for it based on the pace at which the model learns from each modality, which we refer to as the conditional learning speed. We propose an algorithm to balance the conditional learning speeds between modalities during training and demonstrate that it indeed addresses the issue of greedy learning. 1 The proposed algorithm improves the model's generalization on three datasets: Colored MNIST, Model-Net40, and NVIDIA Dynamic Hand Gesture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07af5790-986c-4dad-ae0b-2b2b2ad095c3Cited by top-tier papers51
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei et al.ICCV 2023 · 84 citations
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 65 citations
- Classifier-guided Gradient Modulation for Enhanced Multimodal LearningZirun Guo, Tao Jin, Jingyuan Chen, Zhou ZhaoNeurIPS 2024 · 56 citations
Builds on7
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang et al.ICCV 2021 · 94 citations
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen et al.ACM MM 2021 · 22 citations
- Trusted Multi-View ClassificationZongbo Han, Changqing Zhang, Huazhu Fu, Joey Tianyi ZhouICLR 2021
Related papers
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 16 citations
- Balancing Multimodal Domain Generalization via Gradient Modulation and ProjectionHongzhao Li, Guohao Shen, Shupan Li, Mingliang Xu et al.AAAI 2026 · 1 citation
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang et al.NeurIPS 2025 · 17 citations
- PMR: Prototypical Modal Rebalance for Multimodal LearningYunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang et al.CVPR 2023
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
