SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
Tsu-Jui Fu, Xin Wang, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
摘要
Iterative Language-Based Image Editing (IL-BIE) tasks follow iterative instructions to edit images step by step. Data scarcity is a significant issue for ILBIE as it is challenging to collect large-scale examples of images before and after instruction-based changes. However, humans still accomplish these editing tasks even when presented with an unfamiliar image-instruction pair. Such ability results from counterfactual thinking and the ability to think about alternatives to events that have happened already. In this paper, we introduce a Self-Supervised Counterfactual Reasoning (SSCR) framework that incorporates counterfactual thinking to overcome data scarcity. SSCR allows the model to consider out-ofdistribution instructions paired with previous images. With the help of cross-task consistency (CTC), we train these counterfactual instructions in a self-supervised scenario. Extensive results show that SSCR improves the correctness of ILBIE in terms of both object identity and position, establishing a new state of the art (SOTA) on two IBLIE datasets (i-CLEVR and CoDraw). Even with only 50% of the training data, SSCR achieves a comparable result to using complete data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani 等NeurIPS 2023 · 被引用 462 次
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
- Talk-to-Edit: Fine-Grained Facial Editing via DialogYuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy 等ICCV 2021 · 被引用 162 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- M3L: Language-based Video Editing via Multi-Modal Multi-Level TransformersTsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein 等CVPR 2022 · 被引用 13 次
它引用的顶会 Paper1
相关 Paper
- LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningGaoxiang Cong, Liang Li, Zhenhuan Liu, Yunbin Tu 等ACM MM 2022 · 被引用 11 次
- Multi-Level Counterfactual Contrast for Visual Commonsense ReasoningXi Zhang, Feifei Zhang, Changsheng XuACM MM 2021 · 被引用 22 次
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang 等NeurIPS 2025 · 被引用 71 次
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li 等AAAI 2024 · 被引用 12 次
- Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflectionYucheng Suo, Fan Ma, Kaixin Shen, Linchao Zhu 等ICLR 2025
