Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
Siqi Lu, Wanying Xu, Yongbin Zheng, Wenting Luan, Peng Sun, Jianhang Yao
Abstract
Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where the model develops an implicit preference for certain modalities, leading to under-optimization of others. We propose a simple yet efficient method to address this challenge. The central insight of our work is that the dominance relationship between modalities can be effectively discerned and quantified in the frequency domain. To leverage this principle, we first introduce a Frequency Ratio Metric (FRM) to quantify modality preference by analyzing features in the frequency domain. Guided by FRM, we propose a Multimodal Weight Allocation Module, a plug-and-play component that dynamically rebalances the contribution of each branch during training, thereby promoting a more holistic learning paradigm. Extensive experiments demonstrate that MWAM can be seamlessly integrated into diverse architectural backbones, such as those based on CNNs and ViTs. Furthermore, MWAM delivers consistent performance gains across a wide range of tasks and modality combinations. This advancement extends beyond merely optimizing the performance of the base model; it also yields further improvements to state-of-the-art methods designed to address the missing modality problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Implicit Neural Representation for Cooperative Low-light Image EnhancementShuzhou Yang, Moxuan Ding, Yanmin Wu, Zihan Li et al.ICCV 2023 · 224 citations
- RFNet: Region-aware Fusion Network for Incomplete Multi-modal Brain Tumor SegmentationYuhang Ding, Xin Yu, Yi YangICCV 2021 · 160 citations
- M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing ModalitiesHong Liu, Dong Wei, Donghuan Lu, Jinghan Sun et al.AAAI 2023 · 101 citations
- Boosting Multi-modal Model Performance with Adaptive Gradient ModulationHong Li, Xingyu Li, Pengbo Hu, Yinuo Lei et al.ICCV 2023 · 84 citations
Related papers
- BALM: A Model-Agnostic Framework for Balanced Multimodal Learning under Imbalanced Missing RatesPhuong-Anh Nguyen, Tien Anh Pham, Duc-Trong Le, Cam-Van Thi NguyenCVPR 2026 · 1 citation
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Adaptive Re-calibration Learning for Balanced Multimodal Intention RecognitionQu Yang, Xiyang Li, Fu Lin, Mang YeNeurIPS 2025 · 2 citations
- Reducing Unimodal Bias in Multi-Modal Semantic Segmentation With Multi-Scale Functional Entropy RegularizationXu Zheng, Yuanhuiyi Lyu, Lutao Jiang, Danda Pani Paudel et al.ICCV 2025 · 2 citations
- Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal LearningHossein Rajoli Nowdeh, Jie Ji, Xiaolong Ma, Fatemeh AfghahNeurIPS 2025 · 3 citations
