StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single Video
Lizhen Wang, Xiaochen Zhao, Jingxiang Sun, Yuxiang Zhang, Hongwen Zhang, Tao Yu, Yebin Liu
Abstract
Face reenactment methods attempt to restore and re-animate portrait videos as realistically as possible. Existing methods face a dilemma in quality versus controllability: 2D GAN-based methods achieve higher image quality but suffer in fine-grained control of facial attributes compared with 3D counterparts. In this work, we propose StyleAvatar, a real-time photo-realistic portrait avatar reconstruction method using StyleGAN-based networks, which can generate high-fidelity portrait avatars with faithful expression control. We expand the capabilities of StyleGAN by introducing a compositional representation and a sliding window augmentation method, which enable faster convergence and improve translation generalization. Specifically, we divide the portrait scenes into three parts for adaptive adjustments: facial region, non-facial foreground region, and the background. Besides, our network leverages the best of UNet, StyleGAN and time coding for video learning, which enables high-quality video generation. Furthermore, a sliding window augmentation method together with a pre-training strategy are proposed to improve translation generalization and training performance, respectively. The proposed network can converge within two hours while ensuring high image quality and a forward rendering time of only 20 milliseconds. Furthermore, we propose a real-time live system, which further pushes research into applications. Results and experiments demonstrate the superiority of our method in terms of image quality, full portrait video generation, and real-time re-animation compared to existing facial reenactment methods. Training and inference code for this paper are at https://github.com/LizhenWangT/StyleAvatar.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a8d13a2-0aec-40f5-a096-e00bb33925c9Cited by top-tier papers24
- AvatarReX: Real-time Expressive Full-body AvatarsZerong Zheng, Xiaochen Zhao, Hongwen Zhang, Boning Liu et al.SIGGRAPH 2023 · 80 citations
- AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and PoseHuichao Zhang, Bowen Chen, Hao Yang, Liao Qu et al.AAAI 2024 · 73 citations
- AvatarMAV: Fast 3D Head Avatar Reconstruction Using Motion-Aware Neural VoxelsYuelang Xu, Lizhen Wang, Xiaochen Zhao, Hongwen Zhang et al.SIGGRAPH 2023 · 69 citations
- LatentAvatar: Learning Latent Expression Code for Expressive Neural Head AvatarYuelang Xu, Hongwen Zhang, Lizhen Wang, Xiaochen Zhao et al.SIGGRAPH 2023 · 40 citations
- LayGA: Layered Gaussian Avatars for Animatable Clothing TransferSiyou Lin, Zhe Li, Zhaoqi Su, Zerong Zheng et al.SIGGRAPH 2024 · 27 citations
Builds on27
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- ReStyle: A Residual-Based StyleGAN Encoder via Iterative RefinementYuval Alaluf, Or Patashnik, Daniel Cohen-OrICCV 2021 · 377 citations
Related papers
- Unsupervised Facial Performance Editing via Vector-Quantized StyleGAN RepresentationsBerkay Kicanaoglu, Pablo Garrido, Gaurav BharajICCV 2023 · 2 citations
- Normalized Avatar Synthesis Using StyleGAN and Perceptual RefinementHuiwen Luo, Koki Nagano, Han-Wei Kung, Qingguo Xu et al.CVPR 2021
- StyleRig: Rigging StyleGAN for 3D Control Over Portrait ImagesAyush Tewari, Mohamed A. Elgharib, Gaurav Bharaj, Florian Bernard et al.CVPR 2020
- AgileGAN: stylizing portraits by inversion-consistent transfer learningGuoxian Song, Linjie Luo, Jing Liu, Wan-Chun Ma et al.SIGGRAPH 2021 · 80 citations
- DeepFaceVideoEditing: sketch-based deep editing of face videosFeng-Lin Liu, Shu-Yu Chen, Yu-Kun Lai, Chunpeng Li et al.SIGGRAPH 2022 · 26 citations
