Quantifying and Enhancing Multi-modal Robustness with Modality Preference
Zequn Yang, Yake Wei, Ce Liang, Di Hu
Abstract
Multi-modal models have shown a promising capability to effectively integrate information from various sources, yet meanwhile, they are found vulnerable to pervasive perturbations, such as uni-modal attacks and missing conditions. To counter these perturbations, robust multi-modal representations are highly expected, which are positioned well away from the discriminative multi-modal decision boundary. In this paper, different from conventional empirical studies, we focus on a commonly used joint multi-modal framework and theoretically discover that larger uni-modal representation margins and more reliable integration for modalities are essential components for achieving higher robustness. This discovery can further explain the limitation of multi-modal robustness and the phenomenon that multi-modal models are often vulnerable to attacks on the specific modality. Moreover, our analysis reveals how the widespread issue, that the model has different preferences for modalities, limits the multi-modal robustness by influencing the essential components and could lead to attacks on the specific modality highly effective. Inspired by our theoretical finding, we introduce a training procedure called Certifiable Robust Multi-modal Training (CRMT), which can alleviate this influence from modality preference and explicitly regulate essential components to significantly improve robustness in a certifiable manner. Our method demonstrates substantial improvements in performance and robustness compared with existing methods. Furthermore, our training procedure can be easily extended to enhance other robust training strategies, highlighting its credibility and flexibility. The code is available at https://github.com/GeWu-Lab/Certifiable-Robust-Multi-modal-Training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
- Enhancing Multimodal Cooperation via Sample-Level Modality ValuationYake Wei, Ruoxuan Feng, Zihe Wang, Di HuCVPR 2024 · 24 citations
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language ModelsJean Park, Kuk Jin Jang, Basam Alasaly, Sriharsha Mopidevi et al.AAAI 2025 · 21 citations
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Multimodal Negative LearningBaoquan Gong, Xiyuan Gao, Pengfei Zhu, Qinghua Hu et al.NeurIPS 2025 · 4 citations
Builds on22
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Certified Robustness to Label-Flipping Attacks via Randomized SmoothingElan Rosenfeld, Ezra Winston, Pradeep Ravikumar, J. Zico KolterICML 2020 · 182 citations
Related papers
- Vulnerability-Aware Robust Multimodal Adversarial TrainingJunrui Zhang, Xinyu Zhao, Jie Peng, Chenjie Wang et al.AAAI 2026
- MMCert: Provable Defense Against Adversarial Attacks to Multi-Modal ModelsYanting Wang, Hongye Fu, Wei Zou, Jinyuan JiaCVPR 2024 · 4 citations
- GMML: Gradient-Modulated Robustness for Imbalance-Aware Multimodal LearningZikai Zhang, Xu Zhang, Ziyi Li, Yidong Li et al.ACM MM 2025 · 1 citation
- Is the Modality Gap a Bug or a Feature? A Robustness PerspectiveRhea Chowers, Oshri Naparstek, Udi Barzelay, Yair WeissCVPR 2026 · 4 citations
- Defending Multimodal Fusion Models Against Single-Source AdversariesKarren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa et al.CVPR 2021
