THEMIS: Regulating Textual Inversion for Personalized Concept Censorship
Yutong Wu, Jie Zhang, Florian Kerschbaum, Tianwei Zhang
摘要
—Personalization has become a crucial demand in the Generative AI technology. As the pre-trained generative model ( e . g ., stable diffusion) has fixed and limited capability, it is desirable for users to customize the model to generate output with new or specific concepts. Fine-tuning the pre-trained model is not a promising solution, due to its high requirements of computation resources and data. Instead, the emerging personalization approaches make it feasible to augment the generative model in a lightweight manner. However, this also induces severe threats if such advanced techniques are misused by malicious users, such as spreading fake news or defaming individual reputations. Thus, it is necessary to regulate personalization models ( i . e ., achieve concept censorship ) for their development and advancement. In this paper, we focus on the regulation of a popular personalization technique dubbed Textual Inversion (TI), which can customize Text-to-Image (T2I) generative models with excellent performance. TI crafts the word embedding that contains detailed information about a specific object. Users can easily add the word embedding to their local T2I model, like the public Stable Diffusion (SD) model, to generate personalized images. The advent of TI has brought about a new business model, evidenced by the public platforms for sharing and selling word embeddings ( e . g ., Civitai [1]). Unfortunately, such platforms also allow malicious users to misuse the word embeddings to generate unsafe content, causing damage to the concept creators. We propose T HEMIS to achieve the personalized concept censorship . Its key idea is to leverage the backdoor technique for good by injecting positive backdoors into the TI embeddings
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion ModelsShiqian Zhao, Chong Wang, Yiming Li, Yihao Huang 等NDSS 2026 · 被引用 3 次
- Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive SmoothingLeyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- When and Where Do Data Poisons Attack Textual Inversion?Jeremy Styborski, Mingzhi Lyu, Jiayou Lu, Nupur Kapur 等ICCV 2025 · 被引用 1 次
- Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional DriftGihoon Kim, Hyungjin Park, Taesup KimICLR 2026 · 被引用 1 次
- Compositional Inversion for Stable Diffusion ModelsXulu Zhang, Xiao-Yong Wei, Jinlin Wu, Tianyi Zhang 等AAAI 2024 · 被引用 26 次
- Towards Reliable Verification of Unauthorized Data Usage in Personalized Text-to-Image Diffusion ModelsBoheng Li, Yanhao Wei, Yankai Fu, Zhenting Wang 等S&P 2025
- Circumventing Concept Erasure Methods For Text-To-Image Generative ModelsMinh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal 等ICLR 2024 · 被引用 82 次
