Doubly Abductive Counterfactual Inference for Text-Based Image Editing
Xue Song, Jiequan Cui, Hanwang Zhang, Jingjing Chen, Richang Hong, Yu-Gang Jiang
Abstract
We study text-based image editing (TBIE) of a single image by counterfactual inference because it is an elegant formulation to precisely address the requirement: the edited image should retain the fidelity of the original one. Through the lens of the formulation, we find that the crux of TBIE is that existing techniques hardly achieve a good trade-off between editability and fidelity, mainly due to the overfitting of the single-image fine-tuning. To this end, we propose a Doubly Abductive Counterfactual inference framework (DAC). We first parameterize an exogenous variable as a UNet LoRA, whose abduction can encode all the image details. Second, we abduct another exogenous variable parameterized by a text encoder LoRA, which recovers the lost editability caused by the overfitted first abduction. Thanks to the second abduction, which exclusively encodes the visual transition from post-edit to preedit, its inversion-subtracting the LoRA-effectively reverts pre-edit back to post-edit, thereby accomplishing the edit. Through extensive experiments, our DAC achieves a good trade-off between editability and fidelity. Thus, we can support a wide spectrum of user editing intents, including addition, removal, manipulation, replacement, style transfer, and facial change, which are extensively validated in both qualitative and quantitative evaluations. Codes are in https://github.com/xuesong39/DAC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 794b1513-9019-46f6-85ef-61289c95feecCited by top-tier papers12
- EditInfinity: Image Editing with Binary-Quantized Generative ModelsJiahuan Wang, Yuxin Chen, Jun Yu, Guangming Lu et al.NeurIPS 2025 · 6 citations
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin et al.ICCV 2025 · 2 citations
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image GenerationBailey Trang Nguyen, Parham Saremi, Alan Q. Wang, Fangrui Huang et al.NeurIPS 2025 · 2 citations
- LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer LearningFengyi Fu, Mengqi Huang, Lei Zhang, Zhendong MaoAAAI 2026 · 1 citation
- Diffusion Counterfactual Generation with Semantic AbductionRajat Rasal, Avinash Kori, Fabio De Sousa Ribeiro, Tian Xia et al.ICML 2025
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual AbductionShanshan Huang, Haoxuan Li, Chunyuan Zheng, Mingyuan Ge et al.CVPR 2025
- UIP2P: Unsupervised Instruction-Based Image Editing via Edit Reversibility ConstraintEnis Simsar, Alessio Tonioni, Yongqin Xian, Thomas Hofmann et al.ICCV 2025 · 1 citation
- TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial EditingYanbo Xu, Yueqin Yin, Liming Jiang, Qianyi Wu et al.CVPR 2022 · 53 citations
- MACE: Mass Concept Erasure in Diffusion ModelsShilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu et al.CVPR 2024 · 40 citations
- SINE: SINgle Image Editing with Text-to-Image Diffusion ModelsZhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris N. Metaxas et al.CVPR 2023
