Classifier-guided Gradient Modulation for Enhanced Multimodal Learning
Zirun Guo, Tao Jin, Jingyuan Chen, Zhou Zhao
摘要
Multimodal learning has developed very fast in recent years. However, during the multimodal training process, the model tends to rely on only one modality based on which it could learn faster, thus leading to inadequate use of other modalities. Existing methods to balance the training process always have some limitations on the loss functions, optimizers and the number of modalities and only consider modulating the magnitude of the gradients while ignoring the directions of the gradients. To solve these problems, in this paper, we present a novel method to balance multimodal learning with Classifier-Guided Gradient Modulation (CGGM), considering both the magnitude and directions of the gradients. We conduct extensive experiments on four multimodal datasets: UPMC-Food 101, CMU-MOSI, IEMOCAP and BraTS 2021, covering classification, regression and segmentation tasks. The results show that CGGM outperforms all the baselines and other state-of-the-art methods consistently, demonstrating its effectiveness and versatility. Our code is available at https://github.com/zrguo/CGGM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation LearningChengxuan Qian, Shuo Xing, Li Li, Yue Zhao 等ICLR 2026 · 被引用 42 次
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 被引用 12 次
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian 等CVPR 2026 · 被引用 9 次
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 等CVPR 2026 · 被引用 7 次
- Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment AnalysisMiaosen Luo, Yuncheng Jiang, Sijie MaiACM MM 2025 · 被引用 7 次
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang 等ICML 2022 · 被引用 168 次
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu 等ICML 2023 · 被引用 143 次
- Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksNan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. GerasICML 2022 · 被引用 124 次
相关 Paper
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou 等ACM MM 2023 · 被引用 5 次
- Geometric Gradient Divergence Modulation for Imbalanced Multimodal LearningDisen Hu, Xun Jiang, Zhe Sun, Hao Yang 等ACM MM 2025 · 被引用 1 次
- Reconcile Gradient Modulation for Harmony Multimodal LearningXiyuan Gao, Bing Cao, Baoquan Gong, Pengfei ZhuAAAI 2026
- Boosting Multimodal Learning via Disentangled Gradient LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 被引用 9 次
- PgM: Partitioner Guided Modal Learning FrameworkGuimin Hu, Yi Xin, Lijie Hu, Zhihong Zhu 等ACM MM 2025 · 被引用 1 次
