Lune

CVPR2026Top-tier venue

SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Training

Zhiyi Duan, Xiaoyue Zhang, Tianxing Man

2026Year

Abstract

Recent Vision-Language Models (VLMs) have become increasingly susceptible to jailbreak attacks, where adversarial prompts exploit subtle manipulation to circumvent safety alignment. The diversity and adaptability of such jailbreakers necessitate a defense mechanism with strong generalization capability. However, fine-tuning large-scale VLMs is computationally expensive, and introducing excessive visual or textual defense prompts is impractical for preserving image realism and model usability. To this end, we propose SafeLogo, which tunes a logo-sized visual prompt into a universal shield against diverse jailbreak attacks through micro-regional adversarial training. We are the first to integrate min-max adversarial optimization into visual defense prompt generation. Specifically, in the outer loop, SafeLogo injects compact, bounded perturbations into extremely small image regions (≤ 2% pixel coverage), effectively preserving both visual fidelity and semantic consistency. Meanwhile, overcoming the limitations of prior defenses constrained to a single attack direction or fixed benign supervision, the inner loop dynamically generates and selects the strongest one from a variety of jailbreakers. Extensive experiments on LLaVA-1.5-13B, MiniGPT-4, and Qwen3-VL show that SafeLogo markedly lowers jailbreak ASR on MM-SafetyBench, VLGuard, and FigStep, while preserving benign performance on MM-Vet and MME.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines