Defending Multimodal Fusion Models Against Single-Source Adversaries
Karren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa, J. Zico Kolter
摘要
Beyond achieving high performance across many vision tasks, multimodal models are expected to be robust to single-source faults due to the availability of redundant information between modalities. In this paper, we investigate the robustness of multimodal neural networks against worst-case (i.e., adversarial) perturbations on a single modality. We first show that standard multimodal fusion models are vulnerable to single-source adversaries: an attack on any single modality can overcome the correct information from multiple unperturbed modalities and cause the model to fail. This surprising vulnerability holds across diverse multimodal tasks and necessitates a solution. Motivated by this finding, we propose an adversarially robust fusion strategy that trains the model to compare information coming from all the input sources, detect inconsistencies in the perturbed modality compared to the other modalities, and only allow information from the unperturbed modalities to pass through. Our approach significantly improves on state-of-the-art methods in singlesource robustness, achieving gains of 7.8-25.2% on action recognition, 19.7-48.2% on object detection, and 1.6-6.7% on sentiment analysis, without degrading performance on unperturbed (i.e., clean) data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 被引用 111 次
- On the Limitations of Stochastic Pre-processing DefensesYue Gao, Ilia Shumailov, Kassem Fawaz, Nicolas PapernotNeurIPS 2022 · 被引用 35 次
- HateProof: Are Hateful Meme Detection Systems really Robust?Piush Aggarwal, Pranit Chawla, Mithun Das, Punyajoy Saha 等WWW 2023 · 被引用 15 次
- One Perturbation is Enough: On Generating Universal Adversarial Perturbations Against Vision-Language Pre-Training ModelsHao Fang, Jiawei Kong, Wenbo Yu, Bin Chen 等ICCV 2025 · 被引用 8 次
- MMCert: Provable Defense Against Adversarial Attacks to Multi-Modal ModelsYanting Wang, Hongye Fu, Wei Zou, Jinyuan JiaCVPR 2024 · 被引用 4 次
它引用的顶会 Paper1
相关 Paper
- Fusion Is Not Enough: Single Modal Attacks on Fusion Models for 3D Object DetectionZhiyuan Cheng, Hongjun Choi, Shiwei Feng, James Chenhao Liang 等ICLR 2024 · 被引用 32 次
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine 等CVPR 2022 · 被引用 153 次
- Vulnerability-Aware Robust Multimodal Adversarial TrainingJunrui Zhang, Xinyu Zhao, Jie Peng, Chenjie Wang 等AAAI 2026
- PAIF: Perception-Aware Infrared-Visible Image Fusion for Attack-Tolerant Semantic SegmentationZhu Liu, Jinyuan Liu, Benzhuang Zhang, Long Ma 等ACM MM 2023 · 被引用 47 次
- Towards Good Practices for Missing Modality Robust Action RecognitionSangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho 等AAAI 2023 · 被引用 80 次
