Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
Zirun Guo, Tao Jin, Jingyuan Chen, Zhou Zhao
Abstract
Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with Classifier-Guided Gradient Modulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https://github.com/zrguo/CGGM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation LearningChengxuan Qian, Shuo Xing, Li Li, Yue Zhao et al.ICLR 2026 · 42 citations
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 12 citations
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu et al.CVPR 2026 · 7 citations
- Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment AnalysisMiaosen Luo, Yuncheng Jiang, Sijie MaiACM MM 2025 · 7 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksNan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. GerasICML 2022 · 124 citations
Related papers
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou et al.ACM MM 2023 · 5 citations
- Geometric Gradient Divergence Modulation for Imbalanced Multimodal LearningDisen Hu, Xun Jiang, Zhe Sun, Hao Yang et al.ACM MM 2025 · 1 citation
- Reconcile Gradient Modulation for Harmony Multimodal LearningXiyuan Gao, Bing Cao, Baoquan Gong, Pengfei ZhuAAAI 2026
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 9 citations
- PgM: Partitioner Guided Modal Learning FrameworkGuimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu et al.ACM MM 2025 · 1 citation
