SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Training
Zhiyi Duan, Xiaoyue Zhang, Tianxing Man
Abstract
Recent Vision-Language Models (VLMs) have become increasingly susceptible to jailbreak attacks, where adversarial prompts exploit subtle manipulation to circumvent safety alignment. The diversity and adaptability of such jailbreakers necessitate a defense mechanism with strong generalization capability. However, fine-tuning large-scale VLMs is computationally expensive, and introducing excessive visual or textual defense prompts is impractical for preserving image realism and model usability. To this end, we propose SafeLogo, which tunes a logo-sized visual prompt into a universal shield against diverse jailbreak attacks through micro-regional adversarial training. We are the first to integrate min-max adversarial optimization into visual defense prompt generation. Specifically, in the outer loop, SafeLogo injects compact, bounded perturbations into extremely small image regions (≤ 2% pixel coverage), effectively preserving both visual fidelity and semantic consistency. Meanwhile, overcoming the limitations of prior defenses constrained to a single attack direction or fixed benign supervision, the inner loop dynamically generates and selects the strongest one from a variety of jailbreakers. Extensive experiments on LLaVA-1.5-13B, MiniGPT-4, and Qwen3-VL show that SafeLogo markedly lowers jailbreak ASR on MM-SafetyBench, VLGuard, and FigStep, while preserving benign performance on MM-Vet and MME.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual PromptsYichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang et al.AAAI 2025 · 350 citations
Related papers
- Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention HijackingJingru Li, Wei Ren, Tianqing ZhuACL 2026
- SafetyReminder: Reviving Delayed Safety Awareness of Vision-Language Models to Defend Against Jailbreak AttacksPeiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun et al.AAAI 2026
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 90 citations
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 198 citations
- Towards Robust Multimodal Large Language Models Against Jailbreak AttacksZiyi Yin, Yuanpu Cao, Han Liu, Ting Wang et al.CVPR 2026 · 5 citations
