Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models
Yixin Liu, Ruoxi Chen, Xun Chen, Lichao Sun
摘要
Personalized diffusion models (PDMs) have become prominent for adapting pre-trained text-to-image models to generate images of specific subjects using minimal training data. However, PDMs are susceptible to minor adversarial perturbations, leading to significant degradation when fine-tuned on corrupted datasets. These vulnerabilities are exploited to create protective perturbations that prevent unauthorized image generation. Existing purification methods attempt to red-team the protective perturbation to break the protection but often over-purify images, resulting in information loss. In this work, we conduct an in-depth analysis of the fine-tuning process of PDMs through the lens of shortcut learning. We hypothesize and empirically demonstrate that adversarial perturbations induce a latent-space misalignment between images and their text prompts in the CLIP embedding space. This misalignment causes the model to erroneously associate noisy patterns with unique identifiers during fine-tuning, resulting in poor generalization. Based on these insights, we propose a systematic red-teaming framework that includes data purification and contrastive decoupling learning. We first employ off-the-shelf image restoration techniques to realign images with their original semantic content in latent space. Then, we introduce contrastive decoupling learning with noise tokens to decouple the learning of personalized concepts from spurious noise patterns. Our study not only uncovers shortcut learning vulnerabilities in PDMs but also provides a thorough evaluation framework for developing stronger protection. Our extensive evaluation demonstrates its advantages over existing purification methods and its robustness against adaptive perturbations. Code is available at https://github.com/liuyixin-louis/DiffShortcut . CCS Concepts • Security and privacy → Privacy protections; • Computing methodologies → Computer vision; • Human-centered computing → Empirical studies in HCI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao 等ICML 2022 · 被引用 663 次
相关 Paper
- DDIM Inversion as a Perturbation Amplifier: Breaking Mimicry Protection via Reconstruction Error MinimizationHuming Qiu, Peiyi Chen, Mi Zhang, Geng Hong 等ICML 2026
- Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based CustomizationYisu Liu, Jinyang An, Wanqian Zhang, Dayan Wu 等ACM MM 2024 · 被引用 16 次
- Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion ModelsTae-Young Lee, Juwon Seo, Jong Hwan Ko, Gyeong-Moon ParkNeurIPS 2025 · 被引用 2 次
- Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature OptimizationZiang Xu, Wenbo Yu, Hongyao Yu, Hao Fang 等KDD 2026 · 被引用 2 次
- RECOVER: Reliable Detection of Unauthorized Data Usage in Text-to-Image Diffusion Models via Inversion RobustnessYanhao Wei, Xiaokang Zhao, Boheng Li, Yang Zhang 等ICML 2026
