Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion
Qing-Yuan Jiang, Longfei Huang, Yang Yang
Abstract
Multimodal learning (MML) is significantly constrained by modality imbalance, leading to suboptimal performance in practice. While existing approaches primarily focus on balancing the learning of different modalities to address this issue, they fundamentally overlook the inherent disproportion in model classification ability, which serves as the primary cause of this phenomenon. In this paper, we propose a novel multimodal learning approach to dynamically balance the classification ability of weak and strong modalities by incorporating the principle of boosting. Concretely, we first propose a sustained boosting algorithm in multimodal learning by simultaneously optimizing the classification and residual errors. Subsequently, we introduce an adaptive classifier assignment strategy to dynamically facilitate the classification performance of the weak modality. Furthermore, we theoretically analyze the convergence property of the cross-modal gap function, ensuring the effectiveness of the proposed boosting scheme. To this end, the classification ability of strong and weak modalities is expected to be balanced, thereby mitigating the imbalance issue. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SOTA) multimodal learning baselines. The source code is available at https://github. com/njustkmg/NeurIPS25-AUG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dd30577-93a6-4d70-a863-191fd083afdfCited by top-tier papers2
- Adaptive Debiasing Tsallis Entropy for Test-Time AdaptationXiangyu Wu, Dongming Jiang, Feng Yu, Yueying Tian et al.ICLR 2026 · 2 citations
- MASAM: Multimodal Adaptive Sharpness-Aware Minimization for Heterogeneous Data FusionZijie Chen, Kejing Yin, Wenfang Yao, William Kwok-Wai Cheung et al.ICLR 2026
Builds on20
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
Related papers
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 65 citations
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Multimodal Negative LearningBaoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu et al.NeurIPS 2025 · 4 citations
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou et al.ACM MM 2023 · 5 citations
- PMR: Prototypical Modal Rebalance for Multimodal LearningYunfeng Fan, Wenchao Xu, Haozhao Wang, Junxiao Wang et al.CVPR 2023
