Exploring Multimodal Diffusion Transformers for Enhanced Prompt-Based Image Editing
Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, Jaesik Park
Abstract
Transformer-based diffusion models have recently superseded traditional U-Net architectures, with multimodal diffusion transformers (MM-DiT) emerging as the dominant approach in state-of-the-art models like Stable Diffusion 3 and Flux.1. Previous approaches have relied on unidirectional cross-attention mechanisms, with information flowing from text embeddings to image latents. In contrast, MMDiT introduces a unified attention mechanism that concatenates input projections from both modalities and performs a single full attention operation, allowing bidirectional information flow between text and image branches. This architectural shift presents significant challenges for existing editing techniques. In this paper, we systematically analyze MM-DiT's attention mechanism by decomposing attention matrices into four distinct blocks, revealing their inherent characteristics. Through these analyses, we propose a robust, prompt-based image editing method for MM-DiT that supports global to local edits across various MM-DiT variants, including few-step models. We believe our findings bridge the gap between existing U-Net-based methods and emerging architectures, offering deeper insights into MMDiT's behavioral patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35fbe8d1-87b2-47b8-a5d4-acfef41bd4efCited by top-tier papers6
- Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar PretrainingJunxuan Li, Rawal Khirodkar, Egor Zakharov, Jihyun Lee et al.CVPR 2026 · 3 citations
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion TransformersKanghyun Baek, Jaihyun Lew, Chaehun Shin, Jungbeom Lee et al.ICML 2026 · 1 citation
- A Training-Free Style-Personalization via SVD-Based Feature DecompositionKyoungmin Lee, Jihun Park, Jongmin Gim, Wonhyeok Choi et al.CVPR 2026
- NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight SpacesJiwoo Kim, Swarajh Mehta, Hao-Lun Hsu, Hyunwoo Ryu et al.ICML 2026
- Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion TransformersYuxuan Yao, Yuxuan Chen, Hui Li, Kaihui Cheng et al.ICML 2026
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin et al.ICCV 2025 · 2 citations
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick et al.ICCV 2025 · 9 citations
- DiT4Edit: Diffusion Transformer for Image EditingKunyu Feng, Yue Ma, Bingyuan Wang, Chenyang Qi et al.AAAI 2025 · 92 citations
- Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information FlowsXiang Yang, Feifei Li, Mi Zhang, Geng Hong et al.ICML 2026
- Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion ModelQingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai et al.ICLR 2026 · 40 citations
