Learning by Planning: Language-Guided Global Image Editing
Jing Shi, Ning Xu, Yihang Xu, Trung Bui, Franck Dernoncourt, Chenliang Xu
Abstract
Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in interpretability. To overcome the collective difficulties, we develop a text-to-operation model to map the vague editing language request into a series of editing operations, e.g., change contrast, brightness, and saturation. Each operation is interpretable and differentiable. Furthermore, the only supervision in the task is the target image, which is insufficient for a stable training of sequential decisions. Hence, we propose a novel operation planning algorithm to generate possible editing sequences from the target image as pseudo ground truth. Comparison experiments on the newly collected MA5k-Req dataset and GIER dataset show the advantages of our methods. Code is available at https://github.com/jshi31/T2ONet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingYusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong et al.CVPR 2026 · 63 citations
- Auto-Encoding Morph-Tokens for Multimodal LLMKaihang Pan, Siliang Tang, Juncheng Li, Zhaoyu Fan et al.ICML 2024 · 36 citations
- Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningYuanhao Zhai, Tianyu Luan, David S. Doermann, Junsong YuanICCV 2023 · 35 citations
- Language-Guided Global Image Editing via Cross-Modal Cyclic MechanismWentao Jiang, Ning Xu, Jiayun Wang, Chen Gao et al.ICCV 2021 · 28 citations
Builds on4
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
- Diverse Image Synthesis From Semantic Layouts via Conditional IMLEKe Li, Tianhao Zhang, Jitendra MalikICCV 2019 · 102 citations
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
Related papers
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 3 citations
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang et al.ACM MM 2021 · 28 citations
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 3 citations
- DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksMing Tao, Bing-Kun Bao, Hao Tang, Fei Wu et al.AAAI 2023 · 19 citations
- ZONE: Zero-Shot Instruction-Guided Local EditingShanglin Li, Bohan Zeng, Yutang Feng, Sicheng Gao et al.CVPR 2024
