Minimal, Local, and Robust: Embedding-Only Edits for Implicit Bias in T2I Models
Feng He, Chao Zhang, Zhixue Zhao
摘要
Implicit assumptions and priors are often necessary in text-to-image generation tasks, especially when textual prompts lack sufficient context.However, these assumptions can sometimes reflect societal biases (e.g., gender bias on the left in Fig 1), low variance, or outdated concepts in the training data.We present Embedding-only Editing (EMBEDIT), a method designed to efficiently edit implicit assumptions and priors in the text-to-image model without affecting unrelated objects or degrading overall performance.Given a "source" prompt (e.g., "nurse") that elicits an assumption (e.g., a female nurse) and a "destination" prompt or distribution (e.g.equal gender chance), EMBEDIT only fine-tunes the word token embedding (WTE) of the target object (i.e.token "nurse"'s WTE).Our method prevents unintended effects on other objects in the model's knowledge base, as the WTEs for unrelated objects and the model weights remain unchanged.Further, our method can be applied to any text-to-image model with a text encoder.It is highly efficient, modifying only 768, 2048, and 4864 parameters for Stable Diffusion 1.4, Stable Diffusion XL, and FLUX, respectively, matching each model's WTE dimension.Additionally, changes could be easily reversed by restoring the original WTE layers.The results show that EMBE-DIT outperforms previous methods in various models, tasks, and editing scenarios (both single and sequential multiple edits), achieving at least a 6.01% improvement (from 87.17% to 93.18%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Erasing Concepts from Diffusion ModelsRohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, David BauICCV 2023 · 被引用 536 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
相关 Paper
- Editing Implicit Assumptions in Text-to-Image Diffusion ModelsHadas Orgad, Bahjat Kawar, Yonatan BelinkovICCV 2023 · 被引用 130 次
- Textualize Visual Prompt for Image Editing via Diffusion BridgePengcheng Xu, Qingnan Fan, Fei Kou, Shuai Qin 等AAAI 2025 · 被引用 4 次
- Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow ModelsVladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas, Tomer MichaeliICCV 2025 · 被引用 30 次
- Fair Text-to-Image Diffusion via Fair MappingJia Li, Lijie Hu, Jingfeng Zhang, Tianhang Zheng 等AAAI 2025 · 被引用 36 次
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion ModelsByeonghu Na, Minsang Park, Gyuwon Sim, Donghyeok Shin 等NeurIPS 2025 · 被引用 8 次
