THEMIS: Regulating Textual Inversion for Personalized Concept Censorship
Yutong Wu, Jie Zhang, Florian Kerschbaum, Tianwei Zhang
Abstract
—Personalization has become a crucial demand in the Generative AI technology. As the pre-trained generative model ( e . g ., stable diffusion) has fixed and limited capability, it is desirable for users to customize the model to generate output with new or specific concepts. Fine-tuning the pre-trained model is not a promising solution, due to its high requirements of computation resources and data. Instead, the emerging personalization approaches make it feasible to augment the generative model in a lightweight manner. However, this also induces severe threats if such advanced techniques are misused by malicious users, such as spreading fake news or defaming individual reputations. Thus, it is necessary to regulate personalization models ( i . e ., achieve concept censorship ) for their development and advancement. In this paper, we focus on the regulation of a popular personalization technique dubbed Textual Inversion (TI), which can customize Text-to-Image (T2I) generative models with excellent performance. TI crafts the word embedding that contains detailed information about a specific object. Users can easily add the word embedding to their local T2I model, like the public Stable Diffusion (SD) model, to generate personalized images. The advent of TI has brought about a new business model, evidenced by the public platforms for sharing and selling word embeddings ( e . g ., Civitai [1]). Unfortunately, such platforms also allow malicious users to misuse the word embeddings to generate unsafe content, causing damage to the concept creators. We propose T HEMIS to achieve the personalized concept censorship . Its key idea is to leverage the backdoor technique for good by injecting positive backdoors into the TI embeddings
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion ModelsShiqian Zhao, Chong Wang, Yiming Li, Yihao Huang et al.NDSS 2026 · 3 citations
- Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive SmoothingLeyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu et al.ICML 2026 · 1 citation
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- When and Where Do Data Poisons Attack Textual Inversion?Jeremy Styborski, Mingzhi Lyu, Jiayou Lu, Nupur Kapur et al.ICCV 2025 · 1 citation
- Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional DriftGihoon Kim, Hyungjin Park, Taesup KimICLR 2026 · 1 citation
- Compositional Inversion for Stable Diffusion ModelsXulu Zhang, Xiao-Yong Wei, Jinlin Wu, Tianyi Zhang et al.AAAI 2024 · 26 citations
- Towards Reliable Verification of Unauthorized Data Usage in Personalized Text-to-Image Diffusion ModelsBoheng Li, Yanhao Wei, Yankai Fu, Zhenting Wang et al.S&P 2025
- Circumventing Concept Erasure Methods For Text-To-Image Generative ModelsMinh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal et al.ICLR 2024 · 82 citations
