Boosting Multimodal Learning via Disentangled Gradient Learning
Shicai Wei, Chunbo Luo, Yang Luo
Abstract
Multimodal learning often encounters the under-optimized problem and may have worse performance than unimodal learning. Existing methods attribute this problem to the imbalanced learning between modalities and rebalance them through gradient modulation. However, they fail to explain why the dominant modality in multimodal models also underperforms that in unimodal learning. In this work, we reveal the optimization conflict between the modality encoder and modality fusion module in multimodal models. Specifically, we prove that the cross-modal fusion in multimodal models decreases the gradient passed back to each modality encoder compared with unimodal models. Consequently, the performance of each modality in the multimodal model is inferior to that in the unimodal model. To this end, we propose a disentangled gradient learning (DGL) framework to decouple the optimization of the modality encoder and modality fusion module in the multimodal model. DGL truncates the gradient back-propagated from the multimodal loss to the modality encoder and replaces it with the gradient from unimodal loss. Besides, DGL removes the gradient back-propagated from the unimodal loss to the modality fusion module. This helps eliminate the gradient interference between the modality encoder and modality fusion module while ensuring their respective optimization processes. Finally, extensive experiments on multiple types of modalities, tasks, and frameworks with dense cross-modal interaction demonstrate the effectiveness and versatility of the proposed DGL. Code is available at https://github.com/shicaiwei123/ICCV2025-GDL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be6db7d7-6645-4f46-9f42-afced0a83e95Cited by top-tier papers6
- CLCR: Cross-Level Semantic Collaborative Representation for Multimodal LearningChunlei Meng, Guanhong Huang, Rong Fu, Runmin Jian et al.CVPR 2026 · 9 citations
- Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents CollaborationChunlei Meng, Pengbin Feng, Rong Fu, Hoi Leong Lee et al.ICML 2026 · 2 citations
- Multimodal Nested Learning for Decoupled and Coordinated OptimizationYanglin Feng, Yang Qin, Dezhong Peng, Rui Wang et al.ICML 2026
- No Data? No Problem: Robust Vision-Tabular Learning with Missing ValuesMarta Hasny, Laura Daza, Keno Bressem, Maxime Di Folco et al.ICML 2026
- Multimodal Learning on Low-Quality Data with Conformal Predictive Self-CalibrationXun Jiang, Yufan Gu, Disen Hu, Yuqing Hou et al.CVPR 2026
Builds on13
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic SegmentationJinming Cao, Hanchao Leng, Dani Lischinski, Danny Cohen-Or et al.ICCV 2021 · 186 citations
- RFNet: Region-aware Fusion Network for Incomplete Multi-modal Brain Tumor SegmentationYuhang Ding, Xin Yu, Yi YangICCV 2021 · 160 citations
- CubeMLP: An MLP-based Model for Multimodal Sentiment Analysis and Depression EstimationHao Sun, Hongyi Wang, Jiaqing Liu, Yen-Wei Chen et al.ACM MM 2022 · 158 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
Related papers
- Improving Multimodal Learning via Imbalanced LearningShicai Wei, Chunbo Luo, Yang LuoICCV 2025 · 7 citations
- Detached and Interactive Multimodal LearningYunfeng Fan, Wenchao Xu, Haozhao Wang, Junhong Liu et al.ACM MM 2024 · 4 citations
- Multimodal Representation Learning by Alternating Unimodal AdaptationXiaohui Zhang, Jaehong Yoon, Mohit Bansal, Huaxiu YaoCVPR 2024
- Intra- and Inter-Modal Curriculum for Multimodal LearningYuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan et al.ACM MM 2023 · 28 citations
- G2D: Boosting Multimodal Learning with Gradient-Guided DistillationMohammed Rakib, Arunkumar BagavathiICCV 2025 · 1 citation
