MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, Yinqiang Zheng
Abstract
Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results. For example, generation approaches usually fail to synthesize multiple images of the same objects/characters but with different views or poses. Meanwhile, existing editing methods either fail to achieve effective complex nonrigid editing while maintaining the overall textures and identity, or require time-consuming fine-tuning to capture the image-specific appearance. In this paper, we develop MasaCtrl, a tuning-free method to achieve consistent image generation and complex non-rigid image editing simultaneously. Specifically, MasaCtrl converts existing self-attention in diffusion models into mutual self-attention, so that it can query correlated local contents and textures from source images for consistency. To further alleviate the query confusion between foreground and background, we propose a mask-guided mutual self-attention strategy, where the mask can be easily extracted from the cross-attention maps. Extensive experiments show that the proposed MasaCtrl can produce impressive results in both consistent image generation and complex non-rigid real image editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a9be7bff-8425-4503-881a-ed540c2d8a77Cited by top-tier papers385
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu et al.ICML 2024 · 231 citations
- SEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionXinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang et al.ICLR 2024 · 226 citations
- DragonDiffusion: Enabling Drag-style Manipulation on Diffusion ModelsChong Mou, Xintao Wang, Jiechong Song, Ying Shan et al.ICLR 2024 · 223 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Multi-party Collaborative Attention Control for Image CustomizationHan Yang, Chuanguang Yang, Qiuli Wang, Zhulin An et al.CVPR 2025
- Tuning-Free Inversion-Enhanced Control for Consistent Image EditingXiaoyue Duan, Shuhao Cui, Guoliang Kang, Baochang Zhang et al.AAAI 2024 · 13 citations
- Move and Act: Enhanced Object Manipulation and Background Integrity for Image EditingPengfei Jiang, Mingbao Lin, Fei ChaoAAAI 2025 · 5 citations
- Training-Free Text-Guided Color Editing with Multi-Modal Diffusion TransformerZixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou et al.ICLR 2026 · 6 citations
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt et al.NeurIPS 2023 · 108 citations
