A Latent Transformer for Disentangled Face Editing in Images and Videos
Xu Yao, Alasdair Newson, Yann Gousseau, Pierre Hellier
Abstract
High quality facial image editing is a challenging problem in the movie post-production industry, requiring a high degree of control and identity preservation. Previous works that attempt to tackle this problem may suffer from the entanglement of facial attributes and the loss of the person’s identity. Furthermore, many algorithms are limited to a certain task. To tackle these limitations, we propose to edit facial attributes via the latent space of a StyleGAN generator, by training a dedicated latent transformation network and incorporating explicit disentanglement and identity preservation terms in the loss function. We further introduce a pipeline to generalize our face editing to videos. Our model achieves a disentangled, controllable, and identity-preserving facial attribute editing, even in the challenging case of real (i.e., non-synthetic) images and videos. We conduct extensive experiments on image and video datasets and show that our model outperforms other state-of-the-art methods in visual quality and quantitative evaluation. Source codes are available at https://github.com/InterDigitalInc/latent-transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44f75f6d-4dcb-4cac-9e99-d116d1062540Cited by top-tier papers16
- Style Transformer for Image Inversion and EditingXueqi Hu, Qiusheng Huang, Zhengyi Shi, Siyuan Li et al.CVPR 2022 · 58 citations
- Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language ModelZipeng Xu, Tianwei Lin, Hao Tang, Fu Li et al.CVPR 2022 · 38 citations
- StyleT2I: Toward Compositional and High-Fidelity Text-to-Image SynthesisZhiheng Li, Martin Renqiang Min, Kai Li, Chenliang XuCVPR 2022 · 38 citations
- RIGID: Recurrent GAN Inversion and Editing of Real Face VideosYangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Ping LuoICCV 2023 · 14 citations
- Adaptive Nonlinear Latent Transformation for Conditional Face EditingZhizhong Huang, Siteng Ma, Junping Zhang, Hongming ShanICCV 2023 · 13 citations
Builds on11
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Swapping Autoencoder for Deep Image ManipulationTaesung Park, Jun-Yan Zhu, Oliver Wang, Jingwan Lu et al.NeurIPS 2020 · 376 citations
- Editing in Style: Uncovering the Local Semantics of GANsEdo Collins, Raja Bala, Bob Price, Sabine SüsstrunkCVPR 2020
- Encoding in Style: A StyleGAN Encoder for Image-to-Image TranslationElad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan et al.CVPR 2021
Related papers
- Conceptual and Hierarchical Latent Space Decomposition for Face EditingSavas Özkan, Mete Özay, Tom RobinsonICCV 2023 · 3 citations
- SDGAN: Disentangling Semantic Manipulation for Facial Attribute EditingWenmin Huang, Weiqi Luo, Jiwu Huang, Xiaochun CaoAAAI 2024 · 20 citations
- TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial EditingYanbo Xu, Yueqin Yin, Liming Jiang, Qianyi Wu et al.CVPR 2022 · 53 citations
- Multi-Directional Subspace Editing in Style-SpaceChen NavehICCV 2023 · 4 citations
- Retrieve in Style: Unsupervised Facial Feature Transfer and RetrievalMin Jin Chong, Wen-Sheng Chu, Abhishek Kumar, David A. ForsythICCV 2021 · 27 citations
