Visual Representation Learning through Causal Intervention for Controllable Image Editing
Shanshan Huang, Haoxuan Li, Chunyuan Zheng, Lei Wang, Guorui Liao, Zhili Gong, Huayi Yang, Li Liu
Abstract
A key challenge for controllable image editing is that visual attributes with semantic meanings are not always independent, resulting in spurious correlations in model training. However, most existing methods ignore such issues, leading to biased causal visual representation learning and unintended changes to unrelated regions or attributes in the edited images. To bridge this gap, we propose a diffusion-based causal visual representation learning framework called CIDiffuser to capture causal representations of visual attributes based on structural causal models to address the spurious correlation. Specifically, we first decompose the image representation into a high-level semantic representation for core attributes of the image and a low-level stochastic representation for other random or less structured aspects, with the former extracted by a semantic encoder and the latter derived via a stochastic encoder. We then introduce a causal effect learning module to capture the direct causal effect, that is, the difference of potential outcomes before and after intervening on the visual attributes. In addition, a diffusion-based learning strategy is designed to optimize the representation learning process. Empirical evaluations on two benchmark datasets demonstrate that our approach significantly outperforms state-ofthe-art methods, enabling highly controllable image editing by modifying learned visual representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62773bc3-9dec-4fc0-8ceb-8512d4aa908dCited by top-tier papers3
- Unveiling Extraneous Sampling Bias with Data Missing-Not-At-RandomChunyuan Zheng, Haocheng Yang, Haoxuan Li, Mengyue YangNeurIPS 2025 · 15 citations
- Addressing Correlated Latent Exogenous Variables in Debiased Recommender SystemsShuqiang Zhang, Yuchao Zhang, Jinkun Chen, Haochen SuiKDD 2025 · 4 citations
- Mitigating Data Imbalance in Time Series Classification Based on Counterfactual Minority Samples AugmentationLei Wang, Shanshan Huang, Chunyuan Zheng, Jun Liao et al.KDD 2025 · 1 citation
Builds on20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 330 citations
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 276 citations
- StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion ModelsZhizhong Wang, Lei Zhao, Wei XingICCV 2023 · 219 citations
Related papers
- Diffusion Counterfactual Generation with Semantic AbductionRajat Rasal, Avinash Kori, Fabio De Sousa Ribeiro, Tian Xia et al.ICML 2025
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual GenerationLei Tong, Zhihua Liu, Chaochao Lu, Dino Oglic et al.ICML 2026 · 2 citations
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min et al.ACM MM 2025 · 1 citation
- Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation EditingPeiran Dong, Bingjie Wang, Song Guo, Junxiao Wang et al.NeurIPS 2024 · 4 citations
- Localizing and Editing Knowledge In Text-to-Image Generative ModelsSamyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi et al.ICLR 2024 · 50 citations
