SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
Aafiya Shamshad Hussain, Gaurav Srivastava, Alvi Md. Ishmam, Zaber Ibn Abdul Hakim, Chris Thomas
Abstract
Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal audio-video-language models. We analyze six complementary attack objectives that target different stages of multimodal processing, including audio encoder representations, cross-modal attention, hidden states, and output likelihoods. Across three state-of-the-art models and multiple benchmarks, we show that audio-only perturbations can induce severe multimodal failures, achieving up to 96% attack success rate. We further show that attacks can be successful at low perceptual distortions (LPIPS ≤ 0.08, SI-SNR ≥ 0) and benefit more from extended optimization than increased data scale. Transferability across models and encoders remains limited, while speech recognition systems such as Whisper primarily respond to perturbation magnitude, achieving >97% attack success under severe distortion. These results expose a previously overlooked single-modality attack surface in multimodal systems and motivate defenses that enforce cross-modal consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan Shayegani, Yue Dong, Nael B. Abu-GhazalehICLR 2024 · 271 citations
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 111 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsYuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun et al.NeurIPS 2024 · 67 citations
Related papers
- Cross-modal and Cross-medium Adversarial Attack for AudioLiguo Zhang, Zilin Tian, Yunfei Long, Sizhao Li et al.ACM MM 2023 · 1 citation
- Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation ModelsVyas Raina, Rao Ma, Charles McGhee, Kate M. Knill et al.EMNLP 2024 · 5 citations
- MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference OptimizationAshutosh Chaubey, Jiacheng Pang, Mohammad SoleymaniCVPR 2026 · 7 citations
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu et al.USENIX Security 2026
- Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt InjectionMeng Chen, Kun Wang, Li Lu, Jiaheng Zhang et al.S&P 2026 · 3 citations
