On Mechanistic Knowledge Localization in Text-to-Image Generative Models
Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda, Vlad I. Morariu, Nanxuan Zhao, Ryan A. Rossi, Varun Manjunatha, Soheil Feizi
摘要
Identifying layers within text-to-image models which control visual attributes can facilitate efficient model editing through closed-form updates. Recent work, leveraging causal tracing show that early Stable-Diffusion variants confine knowledge primarily to the first layer of the CLIP text-encoder, while it diffuses throughout the UNet. Extending this framework, we observe that for recent models (e.g., SD-XL, Deep-Floyd), causal tracing fails in pinpointing localized knowledge, highlighting challenges in model editing. To address this issue, we introduce the concept of mechanistic localization in text-toimage models, where knowledge about various visual attributes (e.g., "style", "objects", "facts") can be mechanistically localized to a small fraction of layers in the UNet, thus facilitating efficient model editing. We localize knowledge using our method LOCOGEN which measures the direct effect of intermediate layers to output generation by performing interventions in the cross-attention layers of the UNet. We then employ LOCOEDIT, a fast closed-form editing method across popular open-source text-to-image models (including the latest SD-XL) and explore the possibilities of neuron-level model editing. Using mechanistic localization, our work offers a better view of successes and failures in localization-based textto-image model editing. Code will be available at https://github.com/samyadeepbasu/LocoGen .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Understanding Information Storage and Transfer in Multi-Modal Large Language ModelsSamyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi 等NeurIPS 2024 · 被引用 57 次
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana 等NeurIPS 2024 · 被引用 39 次
- Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language ModelsBaao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang 等NeurIPS 2024 · 被引用 15 次
- Unveiling Concept Attribution in Diffusion ModelsNguyen Hung-Quang, Hoang Phan, Khoa D. DoanNeurIPS 2025 · 被引用 13 次
- TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image GenerationVictor Shea-Jay Huang, Le Zhuo, Yi Xin, Zhaokai Wang 等AAAI 2026 · 被引用 10 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Localizing and Editing Knowledge In Text-to-Image Generative ModelsSamyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi 等ICLR 2024 · 被引用 50 次
- Precise Parameter Localization for Textual Generation in Diffusion ModelsLukasz Staniszewski, Bartosz Cywinski, Franziska Boenisch, Kamil Deja 等ICLR 2025
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt 等NeurIPS 2023 · 被引用 108 次
- TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image EditingSherry X. Chen, Yaron Vaxman, Elad Ben Baruch, David Asulin 等CVPR 2024
- Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image ManipulationXiwen Wei, Zhen Xu, Cheng Liu, Si Wu 等CVPR 2023
