Towards Counterfactual Image Manipulation via CLIP
Yingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang, Shijian Lu, Miaomiao Cui, Xuansong Xie, Xian-Sheng Hua, Chunyan Miao
Abstract
Leveraging StyleGAN's expressivity and its disentangled latent codes, existing methods can achieve realistic editing of different visual attributes such as age and gender of facial images. An intriguing yet challenging problem arises: Can generative models achieve counterfactual editing against their learnt priors? Due to the lack of counterfactual samples in natural datasets, we investigate this problem in a text-driven manner with Contrastive-Language-Image-Pretraining (CLIP), which can offer rich semantic knowledge even for various counterfactual concepts. Different from in-domain manipulation, counterfactual manipulation requires more comprehensive exploitation of semantic knowledge encapsulated in CLIP as well as more delicate handling of editing directions for avoiding being stuck in local minimum or undesired editing. To this end, we design a novel contrastive loss that exploits predefined CLIP-space directions to guide the editing toward desired directions from different perspectives. In addition, we design a simple yet effective scheme that explicitly maps CLIP embeddings (of target text) to the latent space and fuses them with latent codes for effective latent code optimization and accurate editing. Extensive experiments show that our design achieves accurate and realistic editing while driving by target texts with various counterfactual concepts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6cd79cd3-62e6-4e16-9992-d904e769d384Cited by top-tier papers9
- CHAIN: Exploring Global-Local Spatio-Temporal Information for Improved Self-Supervised Video HashingRukai Wei, Yu Liu, Jingkuan Song, Heng Cui et al.ACM MM 2023 · 15 citations
- FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion ModelsHao Zhang, Tianyuan Dai, Yanbo Xu, Yu-Wing Tai et al.NeurIPS 2023 · 9 citations
- Target Bias Is All You Need: Zero-Shot Debiasing of Vision-Language Models With Bias CorpusTaeuk Jang, Hoin Jung, Xiaoqian WangICCV 2025 · 5 citations
- SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image SegmentationKe Yan, Qing Cai, Fan Zhang, Ziyan Cao et al.AAAI 2025 · 1 citation
- Style-Editor: Text-driven Object-centric Style EditingJihun Park, Jongmin Gim, Kyoungmin Lee, Seunghun Lee et al.CVPR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
Related papers
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.CVPR 2022 · 94 citations
- CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face ManipulationChenliang Zhou, Fangcheng Zhong, Cengiz ÖztireliSIGGRAPH 2023 · 13 citations
- StyleGAN-NADA: CLIP-guided domain adaptation of image generatorsRinon Gal, Or Patashnik, Haggai Maron, Amit H. Bermano et al.SIGGRAPH 2022 · 501 citations
- CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCan Wang, Menglei Chai, Mingming He, Dongdong Chen et al.CVPR 2022 · 313 citations
- CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit DirectionsRameen Abdal, Peihao Zhu, John Femiani, Niloy J. Mitra et al.SIGGRAPH 2022 · 76 citations
