Lune

CVPR2026顶会

SafeLogo: Turning Your Logos into Jailbreak Shields via Micro-Regional Adversarial Training

Zhiyi Duan, Xiaoyue Zhang, Tianxing Man

出版方
2026年份

摘要

Recent Vision-Language Models (VLMs) have become increasingly susceptible to jailbreak attacks, where adversarial prompts exploit subtle manipulation to circumvent safety alignment. The diversity and adaptability of such jailbreakers necessitate a defense mechanism with strong generalization capability. However, fine-tuning large-scale VLMs is computationally expensive, and introducing excessive visual or textual defense prompts is impractical for preserving image realism and model usability. To this end, we propose SafeLogo, which tunes a logo-sized visual prompt into a universal shield against diverse jailbreak attacks through micro-regional adversarial training. We are the first to integrate min-max adversarial optimization into visual defense prompt generation. Specifically, in the outer loop, SafeLogo injects compact, bounded perturbations into extremely small image regions (≤ 2% pixel coverage), effectively preserving both visual fidelity and semantic consistency. Meanwhile, overcoming the limitations of prior defenses constrained to a single attack direction or fixed benign supervision, the inner loop dynamically generates and selects the strongest one from a variety of jailbreakers. Extensive experiments on LLaVA-1.5-13B, MiniGPT-4, and Qwen3-VL show that SafeLogo markedly lowers jailbreak ASR on MM-SafetyBench, VLGuard, and FigStep, while preserving benign performance on MM-Vet and MME.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖