An Item Is Worth a Prompt: Versatile Image Editing with Disentangled Control
Aosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong, Kaicheng Zhou, Xiao Zhang, Rex Ying, Leandros Tassiulas
Abstract
Building on the success of text-to-image diffusion models (DPMs), image editing has emerged as a crucial application for enabling human interaction with AI-generated content. Among various editing techniques, prompt-based editing has garnered significant attention for its capacity to simplify semantic control. However, because diffusion models are typically pretrained on descriptive text captions, directly modifying words in text prompts often results in entirely different generated images, which undermines the objectives of image editing. Conversely, existing editing methods often employ spatial masks to maintain the integrity of unedited regions, but these are frequently disregarded by DPMs, leading to disharmonious editing outcomes. To address these two challenges, we propose a method that disentangles the comprehensive image-prompt interaction into multiple item-prompt interactions, with each item associated with a uniquely learned prompt. The resulting framework, named D-Edit, leverages pretrained diffusion models with disentangled cross-attention layers and employs a two-step optimization process to establish item-prompt associations. This approach allows for versatile image editing by enabling targeted manipulations of specific items through their corresponding prompts. We demonstrate state-of-the-art results in four types of editing operations including image-based, text-based, mask-based editing, and item removal, covering most types of editing applications, all within a single unified framework. Notably, D-Edit is the first framework that can (1) achieve item editing through mask editing and (2) combine image and text-based editing. We demonstrate the quality and versatility of the editing results for a diverse collection of images through both qualitative and quantitative evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu et al.CVPR 2026 · 5 citations
- Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion TransferQingyu Shi, Jianzong Wu, Jinbin Bai, Jiangning Zhang et al.ICCV 2025 · 1 citation
- Pose-Star: Anatomy-Aware Editing for Open-World Fashion ImagesYuran Dong, Mang YeICCV 2025 · 1 citation
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 102 citations
- MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted GuidanceQi Mao, Lan Chen, Yuchao Gu, Zhen Fang et al.ACM MM 2024 · 7 citations
- Towards Efficient Diffusion-Based Image Editing with Instant Attention MasksSiyu Zou, Jiji Tang, Yiyi Zhou, Jing He et al.AAAI 2024 · 24 citations
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt et al.NeurIPS 2023 · 108 citations
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
