Unsupervised Facial Performance Editing via Vector-Quantized StyleGAN Representations
Berkay Kicanaoglu, Pablo Garrido, Gaurav Bharaj
Abstract
High-fidelity virtual human avatar applications create a need for photorealistic video face synthesis with controllable semantic editing over facial features. While recent generative neural methods have shown significant progress in portrait video synthesis, intuitive facial control, e.g., of mouth interior and gaze at different levels of details, remains a challenge. In this work, we present a novel face editing framework that combines a 3D face model with StyleGAN vector-quantization to learn multi-level semantic facial control. We show that vector quantization of StyleGAN features unveils richer semantic facial representations, e.g., teeth and pupils, which are difficult to model with 3D tracking priors. Such representations along with 3D tracking can be used as self-supervision to train a generator with control over coarse expressions and finer facial attributes. Learned representations can be combined with user-defined masks to create semantic segmentations that act as custom detail handles for semantic-aware video editing. Our formulation allows video face manipulation with precise local control over facial attributes, such as eyes and teeth, opening up a number of face reenactment and visual expression articulation applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 875cd53f-d76c-48e8-9cbd-2a8951d3ac56Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- ReStyle: A Residual-Based StyleGAN Encoder via Iterative RefinementYuval Alaluf, Or Patashnik, Daniel Cohen-OrICCV 2021 · 377 citations
Related papers
- DeepFaceVideoEditing: sketch-based deep editing of face videosFeng-Lin Liu, Shu-Yu Chen, Yu-Kun Lai, Chunpeng Li et al.SIGGRAPH 2022 · 26 citations
- StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single VideoLizhen Wang, Xiaochen Zhao, Jingxiang Sun, Yuxiang Zhang et al.SIGGRAPH 2023 · 49 citations
- StyleRig: Rigging StyleGAN for 3D Control Over Portrait ImagesAyush Tewari, Mohamed A. Elgharib, Gaurav Bharaj, Florian Bernard et al.CVPR 2020
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- Editing in Style: Uncovering the Local Semantics of GANsEdo Collins, Raja Bala, Bob Price, Sabine SüsstrunkCVPR 2020
