Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
Siyu Zou, Jiji Tang, Yiyi Zhou, Jing He, Chaoyi Zhao, Rongsheng Zhang, Zhipeng Hu, Xiaoshuai Sun
摘要
Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or off-line processing, greatly reducing their efficiency. In this paper, we propose a novel and efficient image editing method for Text-to-Image (T2I) diffusion models, termed Instant Diffusion Editing (InstDiffEdit). In particular, InstDiffEdit aims to employ the cross-modal attention ability of existing diffusion models to achieve instant mask guidance during the diffusion steps. To reduce the noise of attention maps and realize the full automatics, we equip InstDiffEdit with a training-free refinement scheme to adaptively aggregate the attention distributions for the automatic yet accurate mask generation. Meanwhile, to supplement the existing evaluations of DIE, we propose a new benchmark called Editing-Mask to examine the mask accuracy and local editing ability of existing methods. To validate InstDiffEdit, we also conduct extensive experiments on ImageNet and Imagen, and compare it with a bunch of the SOTA methods. The experimental results show that InstDiffEdit not only outperforms the SOTA methods in both image quality and editing results, but also has a much faster inference speed, i.e., +5 to +6 times. Our code available at https://anonymous.4open.science/r/InstDiffEdit-C306
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- VINCIE: Unlocking In-context Image Editing from VideoLeigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao 等ICLR 2026 · 被引用 18 次
- Addressing Text Embedding Leakage in Diffusion-Based Image EditingSunung Mun, Jinhwan Nam, Sunghyun Cho, Jungseul OkICCV 2025 · 被引用 10 次
- Click2Mask: Local Editing with Dynamic Mask GenerationOmer Regev, Omri Avrahami, Dani LischinskiAAAI 2025 · 被引用 3 次
- Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-trainingPeng Sun, Jun XIE, Tao LinCVPR 2026 · 被引用 1 次
- FashionTailor: Controllable Clothing Editing for Human Images with Appearance PreservingJie Hou, Jianghong Ma, Xiangyu Mu, Haijun Zhang 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionTrong-Tung Nguyen, Quang Nguyen, Khoi Nguyen, Anh Tuan Tran 等CVPR 2025
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- An Item Is Worth a Prompt: Versatile Image Editing with Disentangled ControlAosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong 等AAAI 2025 · 被引用 9 次
- Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image EditingBingyan Liu, Chengyu Wang, Tingfeng Cao, Kui Jia 等CVPR 2024
- Inversion-Free Image Editing with Language-Guided Diffusion ModelsSihan Xu, Yidong Huang, Jiayi Pan, Ziqiao Ma 等CVPR 2024 · 被引用 12 次
