Counterfactual Image Editing with Disentangled Causal Latent Space
Yushu Pan, Elias Bareinboim
Abstract
The process of editing an image can be naturally modeled as evaluating a counter-factual query: “What would an image look like if a particular feature had changed?” While recent advances in text-guided image editing leverage powerful pre-trained models to produce visually appealing images, they often lack counterfactual consistency – ignoring how features are causally related and how changing one may affect others. In contrast, existing causal-based editing approaches offer solid theoretical foundations and perform well in specific settings, but remain limited in scalability and often rely on labeled data. In this work, we aim to bridge the gap between causal editing and large-scale text-to-image generation through two main contributions. First, we introduce Backdoor Disentangled Causal Latent Space (BD-CLS), a new class of latent spaces that allows for the encoding of causal inductive biases. One desirable property of this latent space is that, even under weak supervision, it can be shown to exhibit counterfactual consistency. Second, and building on this result, we develop BD-CLS-Edit, an algorithm capable of learning a BD-CLS from a (non-causal) pre-trained Stable Diffusion model. This enables counterfactual image editing without retraining. Our method ensures that edits respect the causal relationships among features, even when some features are unlabeled or unprompted and the original latent space is oblivious to the environment’s underlying cause-and-effect relationships.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f90ffde9-0334-445e-b4e9-620fb3035132Cited by top-tier papers2
- Diverse Text-to-Image Generation via Contrastive Noise OptimizationByungjun Kim, Soobin Um, Jong Chul YeICLR 2026 · 11 citations
- Relational Structural Causal ModelsAdiba Ejaz, Elias BareinboimICML 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min et al.ACM MM 2025 · 1 citation
- Counterfactual Image EditingYushu Pan, Elias BareinboimICML 2024 · 19 citations
- Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data PoisoningShengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu et al.ACM MM 2023 · 46 citations
- Visual Representation Learning through Causal Intervention for Controllable Image EditingShanshan Huang, Haoxuan Li, Chunyuan Zheng, Lei Wang et al.CVPR 2025
- Localizing and Editing Knowledge In Text-to-Image Generative ModelsSamyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi et al.ICLR 2024 · 50 citations
