Lune

ACM MM2025Top-tier venue

AdvPainting: Clean-text Jailbreaking Against Inpainting Models

Bingqian Zhou, Zhihao Wu, Yushi Cheng, Wenyuan Xu

2025Year

Abstract

Text-guided inpainting models are widely used for image editing, restoration, and content generation due to their ability to produce high-fidelity results aligned with natural language prompts. However, these models remain vulnerable to jailbreaking attacks, where adversaries manipulate inputs to generate pornographic or violent content. While prior attacks rely on adversarial text prompts, they are increasingly mitigated by advanced text-based safety filters and manual review. In this work, we propose a new attack paradigm that bypasses these defenses by leveraging the image modality alone. Specifically, we inject imperceptible adversarial perturbations into the input image, enabling successful jailbreaks even when paired with clean prompts (e.g., ''a woman''). To achieve this, we address two key challenges: (1) stabilizing the optimization of adversarial perturbations via a novel gradient estimator, and (2) ensuring visual imperceptibility through a diffusion-based perturbation generator. Extensive experiments show that our method successfully compromises the Stable Diffusion Inpainting model-despite its built-in image and text safety checkers-achieving an average attack success rate (ASR) of 85.7%, significantly outperforming baselines (58.7%). Moreover, our attack exhibits strong transferability across models and maintains robustness against common image pre-processing defenses. Warning: Blurred or masked NSFW imagery is contained.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 50770324-e07d-4789-9f1f-1d5ba89fd8d9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines