SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
Tsu-Jui Fu, Xin Wang, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
Abstract
Iterative Language-Based Image Editing (IL-BIE) tasks follow iterative instructions to edit images step by step. Data scarcity is a significant issue for ILBIE as it is challenging to collect large-scale examples of images before and after instruction-based changes. However, humans still accomplish these editing tasks even when presented with an unfamiliar image-instruction pair. Such ability results from counterfactual thinking and the ability to think about alternatives to events that have happened already. In this paper, we introduce a Self-Supervised Counterfactual Reasoning (SSCR) framework that incorporates counterfactual thinking to overcome data scarcity. SSCR allows the model to consider out-ofdistribution instructions paired with previous images. With the help of cross-task consistency (CTC), we train these counterfactual instructions in a self-supervised scenario. Extensive results show that SSCR improves the correctness of ILBIE in terms of both object identity and position, establishing a new state of the art (SOTA) on two IBLIE datasets (i-CLEVR and CoDraw). Even with only 50% of the training data, SSCR achieves a comparable result to using complete data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddb5b85b-8657-4f1c-a9ec-e2317b8adeb4Cited by top-tier papers9
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy et al.ICCV 2021 · 162 citations
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani et al.ICLR 2023 · 70 citations
- M3L: Language-based Video Editing via Multi-Modal Multi-Level TransformersTsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein et al.CVPR 2022 · 13 citations
Builds on1
Related papers
- LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningGaoxiang Cong, Liang Li, Zhenhuan Liu, Yunbin Tu et al.ACM MM 2022 · 11 citations
- Multi-Level Counterfactual Contrast for Visual Commonsense ReasoningXi Zhang, Feifei Zhang, Changsheng XuACM MM 2021 · 22 citations
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li et al.AAAI 2024 · 12 citations
- Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflectionYucheng Suo, Fan Ma, Kaixin Shen, Linchao Zhu et al.ICLR 2025
