Learning by Planning: Language-Guided Global Image Editing
Jing Shi, Ning Xu, Yihang Xu, Trung Bui, Franck Dernoncourt, Chenliang Xu
摘要
Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in interpretability. To overcome the collective difficulties, we develop a text-to-operation model to map the vague editing language request into a series of editing operations, e.g., change contrast, brightness, and saturation. Each operation is interpretable and differentiable. Furthermore, the only supervision in the task is the target image, which is insufficient for a stable training of sequential decisions. Hence, we propose a novel operation planning algorithm to generate possible editing sequences from the target image as pseudo ground truth. Comparison experiments on the newly collected MA5k-Req dataset and GIER dataset show the advantages of our methods. Code is available at https://github.com/jshi31/T2ONet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingYusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong 等CVPR 2026 · 被引用 63 次
- Auto-Encoding Morph-Tokens for Multimodal LLMKaihang Pan, Siliang Tang, Juncheng Li, Zhaoyu Fan 等ICML 2024 · 被引用 36 次
- Towards Generic Image Manipulation Detection with Weakly-Supervised Self-Consistency LearningYuanhao Zhai, Tianyu Luan, David S. Doermann, Junsong YuanICCV 2023 · 被引用 35 次
- Language-Guided Global Image Editing via Cross-Modal Cyclic MechanismWentao Jiang, Ning Xu, Jiayun Wang, Chen Gao 等ICCV 2021 · 被引用 28 次
它引用的顶会 Paper4
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 被引用 317 次
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm 等ICCV 2019 · 被引用 128 次
- Diverse Image Synthesis From Semantic Layouts via Conditional IMLEKe Li, Tianhao Zhang, Jitendra MalikICCV 2019 · 被引用 102 次
- ManiGAN: Text-Guided Image ManipulationBowen Li, Xiaojuan Qi, Thomas Lukasiewicz, Philip H. S. TorrCVPR 2020
相关 Paper
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 被引用 3 次
- Text as Neural Operator: Image Manipulation by Text InstructionTianhao Zhang, Hung-Yu Tseng, Lu Jiang, Weilong Yang 等ACM MM 2021 · 被引用 28 次
- Target-Free Text-Guided Image ManipulationWan-Cyuan Fan, Cheng-Fu Yang, Chiao-An Yang, Yu-Chiang Frank WangAAAI 2023 · 被引用 3 次
- DE-net: Dynamic Text-Guided Image Editing Adversarial NetworksMing Tao, Bing-Kun Bao, Hao Tang, Fei Wu 等AAAI 2023 · 被引用 19 次
- ZONE: Zero-Shot Instruction-Guided Local EditingShanglin Li, Bohan Zeng, Yutang Feng, Sicheng Gao 等CVPR 2024
