DE-net: Dynamic Text-Guided Image Editing Adversarial Networks
Ming Tao, Bing-Kun Bao, Hao Tang, Fei Wu, Longhui Wei, Qi Tian
Abstract
Text-guided image editing models have shown remarkable results. However, there remain two problems. First, they employ fixed manipulation modules for various editing requirements (e.g., color changing, texture changing, content adding and removing), which results in over-editing or insufficient editing. Second, they do not clearly distinguish between text-required and text-irrelevant parts, which leads to inaccurate editing. To solve these limitations, we propose: (i) a Dynamic Editing Block (DEBlock) which composes different editing modules dynamically for various editing requirements. (ii) a Composition Predictor (Comp-Pred) which predicts the composition weights for DEBlock according to the inference on target texts and source images. (iii) a Dynamic text-adaptive Convolution Block (DCBlock) which queries source image features to distinguish text-required parts and text-irrelevant parts. Extensive experiments demonstrate that our DE-Net achieves excellent performance and manipulates source images more correctly and accurately. Code is available at https: //github.com/tobran/DE-Net .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dbe112e-1cfd-49af-9041-9309769a5b8bCited by top-tier papers3
- Alchemist: Parametric Control of Material Properties with Diffusion ModelsPrafull Sharma, Varun Jampani, Yuanzhen Li, Xuhui Jia et al.CVPR 2024 · 16 citations
- GALIP: Generative Adversarial CLIPs for Text-to-Image SynthesisMing Tao, Bing-Kun Bao, Hao Tang, Changsheng XuCVPR 2023
- Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image InpaintingSu Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont-Tuset et al.CVPR 2023
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
Related papers
- Combining Attention with Flow for Person Image SynthesisYurui Ren, Yubo Wu, Thomas H. Li, Shan Liu et al.ACM MM 2021 · 16 citations
- Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentEn Ci, Shanyan Guan, Yanhao Ge, Yilin Zhang et al.ICCV 2025
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang et al.ICCV 2025
- Language-Guided Global Image Editing via Cross-Modal Cyclic MechanismWentao Jiang, Ning Xu, Jiayun Wang, Chen Gao et al.ICCV 2021 · 28 citations
- FlowDC: Flow-Based Decoupling-Decay for Complex Image EditingYilei Jiang, Zhen Wang, Yanghao Wang, Jun Yu et al.CVPR 2026 · 4 citations
