The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails
Shuo Shi, Rui Yin, Naen Xu, Jiahao Chen, Chunyi Zhou, Tianyu Du, Zhihui Fu, Jun Wang, Zhaoxiang Wang, Shouling Ji
Abstract
Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking attacks that produce false negatives, the inverse threat of inducing false positives on benign inputs remains unexplored. We introduce Unsafe Induction Attacks, where adversaries distribute imperceptibly perturbed safe images that trigger guard models to reject legitimate user requests, causing a ''Boy Who Cried Wolf'' effect that degrades service availability and erodes trust. This reveals an availability failure mode in deployed safety filters. To realize this threat under diverse user prompts, we propose Unsafe Semantic Distillation (USD), which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances. Evaluated on four state-of-the-art guard models across realistic user simulation scenarios, USD achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety architectures. WARNING: This paper contains harmful content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e5304d4-82d3-47ca-9f0d-ded00bc52f44Builds on23
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang et al.NeurIPS 2023 · 404 citations
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang et al.AAAI 2025 · 350 citations
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 111 citations
- UltraRE: Enhancing RecEraser for Recommendation Unlearning via Error DecompositionYuyuan Li, Chaochao Chen, Yizhao Zhang, Weiming Liu et al.NeurIPS 2023 · 90 citations
Related papers
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 20 citations
- Robustness of Vision Language Models Against Split-Image Harmful Input AttacksMd Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta MehnazCCS 2026 · 1 citation
- Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine UnlearningYiwei Chen, Yuguang Yao, Yihua Zhang, Bingquan Shen et al.ICLR 2026
- Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI ResponsesDavid Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan et al.ICLR 2025
- Towards Robust Multimodal Large Language Models Against Jailbreak AttacksZiyi Yin, Yuanpu Cao, Han Liu, Ting Wang et al.CVPR 2026 · 5 citations
