Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism
Wentao Jiang, Ning Xu, Jiayun Wang, Chen Gao, Jing Shi, Zhe Lin, Si Liu
Abstract
Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced and insufficient data distribution of real-world datasets and thus fail to understand language requests well. To handle this issue, we propose to create a cycle with our image generator by creating a novel model called Editing Description Network (EDNet) which predicts an editing embedding given a pair of images. Given the cycle, we propose several free augmentation strategies to help our model understand various editing requests given the imbalanced dataset. In addition, two other novel ideas are proposed: an Image-Request Attention (IRA) module which allows our method to edit an image spatial-adaptively when the image requires different editing degree at different regions, as well as a new evaluation metric for this task which is more semantic and reasonable than conventional pixel losses (e.g. L1). Extensive experiments on two benchmark datasets demonstrate the effectiveness of our method over existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da33c039-87d1-46f1-885b-b9922904cbc1Cited by top-tier papers8
- Sound-Guided Semantic Image ManipulationSeung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon et al.CVPR 2022 · 43 citations
- Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningYuanhao Zhai, Tianyu Luan, David S. Doermann, Junsong YuanICCV 2023 · 35 citations
- DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksMing Tao, Bing-Kun Bao, Hao Tang, Fei Wu et al.AAAI 2023 · 19 citations
- SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color EditingJing Shi, Ning Xu, Haitian Zheng, Alex Smith et al.CVPR 2022 · 15 citations
- ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-wise Semantic Alignment and GenerationJianan Wang, Guansong Lu, Hang Xu, Zhenguo Li et al.CVPR 2022 · 15 citations
Builds on4
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
- Learning by Planning: Language-Guided Global Image EditingJing Shi, Ning Xu, Yihang Xu, Trung Bui et al.CVPR 2021
- AdversarialNAS: Adversarial Neural Architecture Search for GANsChen Gao, Yunpeng Chen, Si Liu, Zhenxiong Tan et al.CVPR 2020
Related papers
- Describe, Don't Dictate: Semantic Image Editing with Natural Language IntentEn Ci, Shanyan Guan, Yanhao Ge, Yilin Zhang et al.ICCV 2025
- AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaQifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan et al.CVPR 2025
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image EditingZhiyuan Ma, Guoli Jia, Bowen ZhouAAAI 2024 · 13 citations
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 3 citations
