Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard
Yudong Yang, Xuezhen Zhang, Zhifeng Han, Siyin Wang, Jimin Zhuang, Zengrui Jin, Jing Shao, Guangzhi Sun, Chao Zhang
摘要
Recent progress in LLMs has enabled understanding of audio signals, but has also exposed new safety risks arising from complex audio inputs that are inadequately handled by current safeguards. We introduce SACRED-Bench (Speech-Audio Composition for RED-teaming) to evaluate the robustness of LLMs under complex audio-based attacks. Unlike existing perturbation-based methods that rely on noise optimization or white-box access, SACRED-Bench exploits speech-audio composition to enable effective black-box attacks. SACRED-Bench adopts three composition mechanisms: (a) overlap of harmful and benign speech, (b) mixture of benign speech with harmful non-speech audio, and (c) multi-speaker dialogue. These mechanisms focus on evaluating safety in settings where benign and harmful intents co-occur within a single auditory scene. Moreover, questions in SACRED-Bench are designed to implicitly refer to content in the audio, such that no explicit harmful information appears in the text prompt alone. Experiments demonstrate that even Gemini 2.5 Pro, a state-of-the-art proprietary LLM with safety guardrails fully enabled, still exhibits a 66% attack success rate. To bridge this gap, we propose SALMONN-Guard, the first guard model that jointly inspects speech, audio, and text for safety judgments, reducing the attack success rate to 20%. Our results highlight the need for audio-aware defenses to ensure the safety of multimodal LLMs. The dataset and SALMONN-Guard checkpoints can be found at https://huggingface.co/datasets/ tsinghua-ee/SACRED-Bench .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
相关 Paper
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language ModelsZifan Peng, Yule Liu, Zhen Sun, Mingchen Li 等ICLR 2026 · 被引用 20 次
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language ModelsWeifei Jin, Yuxin Cao, Junjie Su, Minhui Xue 等NeurIPS 2025 · 被引用 9 次
- MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and ModalitiesSahil Verma, Keegan Hines, Jeff A. Bilmes, Charlotte Siska 等EMNLP 2025
- Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language ModelsZirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li 等ACL 2026 · 被引用 20 次
- SPIRIT: Patching Speech Language Models against Jailbreak AttacksAmirbek Djanibekov, Nurdaulet Mukhituly, Kentaro Inui, Hanan Aldarmaki 等EMNLP 2025 · 被引用 3 次
