Visual Representation Learning through Causal Intervention for Controllable Image Editing
Shanshan Huang, Haoxuan Li, Chunyuan Zheng, Lei Wang, Guorui Liao, Zhili Gong, Huayi Yang, Li Liu
摘要
A key challenge for controllable image editing is that visual attributes with semantic meanings are not always independent, resulting in spurious correlations in model training. However, most existing methods ignore such issues, leading to biased causal visual representation learning and unintended changes to unrelated regions or attributes in the edited images. To bridge this gap, we propose a diffusion-based causal visual representation learning framework called CIDiffuser to capture causal representations of visual attributes based on structural causal models to address the spurious correlation. Specifically, we first decompose the image representation into a high-level semantic representation for core attributes of the image and a low-level stochastic representation for other random or less structured aspects, with the former extracted by a semantic encoder and the latter derived via a stochastic encoder. We then introduce a causal effect learning module to capture the direct causal effect, that is, the difference of potential outcomes before and after intervening on the visual attributes. In addition, a diffusion-based learning strategy is designed to optimize the representation learning process. Empirical evaluations on two benchmark datasets demonstrate that our approach significantly outperforms state-ofthe-art methods, enabling highly controllable image editing by modifying learned visual representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Unveiling Extraneous Sampling Bias with Data Missing-Not-At-RandomChunyuan Zheng, Haocheng Yang, Haoxuan Li, Mengyue YangNeurIPS 2025 · 被引用 15 次
- Addressing Correlated Latent Exogenous Variables in Debiased Recommender SystemsShuqiang Zhang, Yuchao Zhang, Jinkun Chen, Haochen SuiKDD 2025 · 被引用 4 次
- Mitigating Data Imbalance in Time Series Classification Based on Counterfactual Minority Samples AugmentationLei Wang, Shanshan Huang, Chunyuan Zheng, Jun Liao 等KDD 2025 · 被引用 1 次
它引用的顶会 Paper20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identificationYongming Rao, Guangyi Chen, Jiwen Lu, Jie ZhouICCV 2021 · 被引用 330 次
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 被引用 276 次
- StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion ModelsZhizhong Wang, Lei Zhao, Wei XingICCV 2023 · 被引用 219 次
相关 Paper
- Diffusion Counterfactual Generation with Semantic AbductionRajat Rasal, Avinash Kori, Fabio De Sousa Ribeiro, Tian Xia 等ICML 2025
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual GenerationLei Tong, Zhihua Liu, Chaochao Lu, Dino Oglic 等ICML 2026 · 被引用 2 次
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min 等ACM MM 2025 · 被引用 1 次
- Towards Safe Concept Transfer of Multi-Modal Diffusion via Causal Representation EditingPeiran Dong, Bingjie Wang, Song Guo, Junxiao Wang 等NeurIPS 2024 · 被引用 4 次
- Localizing and Editing Knowledge In Text-to-Image Generative ModelsSamyadeep Basu, Nanxuan Zhao, Vlad I. Morariu, Soheil Feizi 等ICLR 2024 · 被引用 50 次
