Doubly Abductive Counterfactual Inference for Text-Based Image Editing
Xue Song, Jiequan Cui, Hanwang Zhang, Jingjing Chen, Richang Hong, Yu-Gang Jiang
摘要
We study text-based image editing (TBIE) of a single image by counterfactual inference because it is an elegant formulation to precisely address the requirement: the edited image should retain the fidelity of the original one. Through the lens of the formulation, we find that the crux of TBIE is that existing techniques hardly achieve a good trade-off between editability and fidelity, mainly due to the overfitting of the single-image fine-tuning. To this end, we propose a Doubly Abductive Counterfactual inference framework (DAC). We first parameterize an exogenous variable as a UNet LoRA, whose abduction can encode all the image details. Second, we abduct another exogenous variable parameterized by a text encoder LoRA, which recovers the lost editability caused by the overfitted first abduction. Thanks to the second abduction, which exclusively encodes the visual transition from post-edit to preedit, its inversion-subtracting the LoRA-effectively reverts pre-edit back to post-edit, thereby accomplishing the edit. Through extensive experiments, our DAC achieves a good trade-off between editability and fidelity. Thus, we can support a wide spectrum of user editing intents, including addition, removal, manipulation, replacement, style transfer, and facial change, which are extensively validated in both qualitative and quantitative evaluations. Codes are in https://github.com/xuesong39/DAC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- EditInfinity: Image Editing with Binary-Quantized Generative ModelsJiahuan Wang, Yuxin Chen, Jun Yu, Guangming Lu 等NeurIPS 2025 · 被引用 6 次
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin 等ICCV 2025 · 被引用 2 次
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image GenerationBailey Trang Nguyen, Parham Saremi, Alan Q. Wang, Fangrui Huang 等NeurIPS 2025 · 被引用 2 次
- LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer LearningFengyi Fu, Mengqi Huang, Lei Zhang, Zhendong MaoAAAI 2026 · 被引用 1 次
- Diffusion Counterfactual Generation with Semantic AbductionRajat Rasal, Avinash Kori, Fabio De Sousa Ribeiro, Tian Xia 等ICML 2025
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual AbductionShanshan Huang, Haoxuan Li, Chunyuan Zheng, Mingyuan Ge 等CVPR 2025
- UIP2P: Unsupervised Instruction-Based Image Editing via Edit Reversibility ConstraintEnis Simsar, Alessio Tonioni, Yongqin Xian, Thomas Hofmann 等ICCV 2025 · 被引用 1 次
- TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial EditingYanbo Xu, Yueqin Yin, Liming Jiang, Qianyi Wu 等CVPR 2022 · 被引用 53 次
- MACE: Mass Concept Erasure in Diffusion ModelsShilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu 等CVPR 2024 · 被引用 40 次
- SINE: SINgle Image Editing with Text-to-Image Diffusion ModelsZhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris N. Metaxas 等CVPR 2023
