SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Training
Zhiyi Duan, Xiaoyue Zhang, Tianxing Man
摘要
Recent Vision-Language Models (VLMs) have become increasingly susceptible to jailbreak attacks, where adversarial prompts exploit subtle manipulation to circumvent safety alignment. The diversity and adaptability of such jailbreakers necessitate a defense mechanism with strong generalization capability. However, fine-tuning large-scale VLMs is computationally expensive, and introducing excessive visual or textual defense prompts is impractical for preserving image realism and model usability. To this end, we propose SafeLogo, which tunes a logo-sized visual prompt into a universal shield against diverse jailbreak attacks through micro-regional adversarial training. We are the first to integrate min-max adversarial optimization into visual defense prompt generation. Specifically, in the outer loop, SafeLogo injects compact, bounded perturbations into extremely small image regions (≤ 2% pixel coverage), effectively preserving both visual fidelity and semantic consistency. Meanwhile, overcoming the limitations of prior defenses constrained to a single attack direction or fixed benign supervision, the inner loop dynamically generates and selects the strongest one from a variety of jailbreakers. Extensive experiments on LLaVA-1.5-13B, MiniGPT-4, and Qwen3-VL show that SafeLogo markedly lowers jailbreak ASR on MM-SafetyBench, VLGuard, and FigStep, while preserving benign performance on MM-Vet and MME.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang 等AAAI 2025 · 被引用 350 次
相关 Paper
- Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention HijackingJingru Li, Wei Ren, Tianqing ZhuACL 2026
- SafetyReminder: Reviving Delayed Safety Awareness of Vision-Language Models to Defend Against Jailbreak AttacksPeiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun 等AAAI 2026
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 被引用 90 次
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 被引用 198 次
- Towards Robust Multimodal Large Language Models Against Jailbreak AttacksZiyi Yin, Yuanpu Cao, Han Liu, Ting Wang 等CVPR 2026 · 被引用 5 次
