ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
Alec Helbling, Tuna Han Salih Meral, Benjamin Hoover, Pinar Yanardag, Duen Horng Chau
摘要
Do the rich representations of multi-modal diffusion transformers (DiTs) exhibit unique properties that enhance their interpretability? We introduce CONCEPTATTENTION, a novel method that leverages the expressive power of DiT attention layers to generate high-quality saliency maps that precisely locate textual concepts within images 4 . Without requiring additional training, CONCEPTATTENTION repurposes the parameters of DiT attention layers to produce highly contextualized concept embeddings, contributing the major discovery that performing linear projections in the output space of DiT attention layers yields significantly sharper saliency maps compared to commonly used cross-attention mechanisms. Remarkably, CONCEPTATTENTION even achieves state-of-the-art performance on zeroshot image segmentation benchmarks, outperforming 11 other zero-shot interpretability methods on the ImageNet-Segmentation dataset and on a single-class subset of PascalVOC. Our work contributes the first evidence that the representations of multi-modal DiT models like Flux are highly transferable to vision tasks like segmentation, even outperforming multi-modal foundation models like CLIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow TransformersYusuf Dalva, Hidir Yesiltepe, Pinar YanardagNeurIPS 2025 · 被引用 13 次
- Attention (as Discrete-Time Markov) ChainsYotam Erel, Olaf Dünkel, Rishabh Dabral, Vladislav Golyanik 等NeurIPS 2025 · 被引用 12 次
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed IndividualsDavide Lobba, Fulvio Sanguigni, Bin Ren, Marcella Cornia 等ICLR 2026 · 被引用 7 次
- Localizing Knowledge in Diffusion TransformersArman Zarei, Samyadeep Basu, Keivan Rezaei, Zihao Lin 等NeurIPS 2025 · 被引用 7 次
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned InterventionsAda Görgün, Fawaz Sammani, Nikos Deligiannis, Bernt Schiele 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon 等NeurIPS 2025 · 被引用 6 次
- Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion TransformersYoungjun Jun, Seil Kang, Woojung Han, Seong Jae HwangCVPR 2026 · 被引用 1 次
- Responsible Text-to-Image Diffusion: Interpretable and Linearly Controllable Semantics for Fair and Safe GenerationSayedmoslem Shokrolahi, Jae-Mo Kang, Il-Min KimICML 2026
- Circuit Mechanisms for Spatial Relation Generation in Diffusion TransformersBinxu Wang, Jingxuan Fan, Xu PanCVPR 2026 · 被引用 4 次
- FreeCus: Free Lunch Subject-Driven Customization in Diffusion TransformersYanbing Zhang, Zhe Wang, Qin Zhou, Mengping YangICCV 2025
