Talk-to-Edit: Fine-Grained Facial Editing via Dialog
Yuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy, Ziwei Liu
摘要
Facial editing is an important task in vision and graphics with numerous applications. However, existing works are incapable to deliver a continuous and fine-grained editing mode (e.g., editing a slightly smiling face to a big laughing one) with natural interactions with users. In this work, we propose Talk-to-Edit, an interactive facial editing framework that performs fine-grained attribute manipulation through dialog between the user and the system. Our key insight is to model a continual "semantic field" in the GAN latent space. 1) Unlike previous works that regard the editing as traversing straight lines in the latent space, here the fine-grained editing is formulated as finding a curving trajectory that respects fine-grained attribute landscape on the semantic field. 2) The curvature at each step is location-specific and determined by the input image as well as the users’ language requests. 3) To engage the users in a meaningful dialog, our system generates language feedback by considering both the user request and the current state of the semantic field.We also contribute CelebA-Dialog, a visual-language facial editing dataset to facilitate large-scale study. Specifically, each image has manually annotated fine-grained attribute annotations as well as template-based textual descriptions in natural language. Extensive quantitative and qualitative experiments demonstrate the superiority of our framework in terms of 1) the smoothness of fine-grained editing, 2) the identity/attribute preservation, and 3) the visual photorealism and dialog fluency. Notably, user study validates that our overall system is consistently favored by around 80% of the participants. Our project page is https://www.mmlab-ntu.com/project/talkedit/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- End-to-End Reconstruction-Classification Learning for Face Forgery DetectionJunyi Cao, Chao Ma, Taiping Yao, Shen Chen 等CVPR 2022 · 被引用 327 次
- AvatarCLIP: zero-shot text-driven generation and animation of 3D avatarsFangzhou Hong, Mingyuan Zhang, Liang Pan, Zhongang Cai 等SIGGRAPH 2022 · 被引用 213 次
- Text2Human: text-driven controllable human image generationYuming Jiang, Shuai Yang, Haonan Qiu, Wayne Wu 等SIGGRAPH 2022 · 被引用 140 次
- VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion ModelsSheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 被引用 101 次
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao 等CVPR 2022 · 被引用 94 次
它引用的顶会 Paper16
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 被引用 1,049 次
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 被引用 459 次
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 被引用 421 次
- Enjoy Your Editing: Controllable GANs for Image Editing via Latent Space NavigationPeiye Zhuang, Oluwasanmi Koyejo, Alexander G. SchwingICLR 2021 · 被引用 88 次
- Sequential Attention GAN for Interactive Image EditingYu Cheng, Zhe Gan, Yitong Li, Jingjing Liu 等ACM MM 2020 · 被引用 74 次
相关 Paper
- ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueXing Cui, Zekun Li, Pei Li, Yibo Hu 等EMNLP 2023 · 被引用 4 次
- SDGAN: Disentangling Semantic Manipulation for Facial Attribute EditingWenmin Huang, Weiqi Luo, Jiwu Huang, Xiaochun CaoAAAI 2024 · 被引用 20 次
- TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial EditingYanbo Xu, Yueqin Yin, Liming Jiang, Qianyi Wu 等CVPR 2022 · 被引用 53 次
- EditGAN: High-Precision Semantic Image EditingHuan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim 等NeurIPS 2021 · 被引用 248 次
- Adaptive Nonlinear Latent Transformation for Conditional Face EditingZhizhong Huang, Siteng Ma, Junping Zhang, Hongming ShanICCV 2023 · 被引用 13 次
