CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable and Controllable Text-Guided Face Manipulation
Chenliang Zhou, Fangcheng Zhong, Cengiz Öztireli
摘要
Recently introduced Contrastive Language-Image Pre-Training (CLIP) [Radford et al. 2021] bridges images and text by embedding them into a joint latent space. This opens the door to ample literature that aims to manipulate an input image by providing a textual explanation. However, due to the discrepancy between image and text embeddings in the joint space, using text embeddings as the optimization target often introduces undesired artifacts in the resulting images. Disentanglement, interpretability, and controllability are also hard to guarantee for manipulation. To alleviate these problems, we propose to define corpus subspaces spanned by relevant prompts to capture specific image characteristics. We introduce CLIP projection-augmentation embedding (PAE) as an optimization target to improve the performance of textguided image manipulation. Our method is a simple and general paradigm that can be easily computed and adapted, and smoothly incorporated into any CLIP-based image manipulation algorithm. To demonstrate the effectiveness of our method, we conduct several theoretical and empirical studies. As a case study, we utilize the method for text-guided semantic face editing. We quantitatively and qualitatively demonstrate that PAE facilitates a more disentangled, interpretable, and controllable face image manipulation with state-of-the-art quality and accuracy. Project page: https://chenliang-zhou.github.io/CLIP-PAE/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative ModelsXingchen Zeng, Ziyao Gao, Yilin Ye, Wei ZengCHI 2024 · 被引用 27 次
- Distributional Vision-Language Alignment by Cauchy-Schwarz DivergenceWenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu 等ICLR 2026 · 被引用 9 次
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 被引用 8 次
- ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion MapYilin Ye, Shishi Xiao, Xingchen Zeng, Wei ZengIEEE VIS 2024 · 被引用 8 次
- M3ashy: Multi-Modal Material Synthesis via HyperdiffusionChenliang Zhou, Zheyuan Hu, Alejandro Sztrajman, Yancheng Cai 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
相关 Paper
- Towards Counterfactual Image Manipulation via CLIPYingchen Yu, Fangneng Zhan, Rongliang Wu, Jiahui Zhang 等ACM MM 2022 · 被引用 33 次
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao 等CVPR 2022 · 被引用 94 次
- CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCan Wang, Menglei Chai, Mingming He, Dongdong Chen 等CVPR 2022 · 被引用 313 次
- CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector GraphicsYiren Song, Xuning Shao, Kang Chen, Weidong Zhang 等AAAI 2023 · 被引用 50 次
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li 等CVPR 2022 · 被引用 38 次
