D2Animator: Dual Distillation of StyleGAN For High-Resolution Face Animation
Zhuo Chen, Chaoyue Wang, Haimei Zhao, Bo Yuan, Xiu Li
Abstract
The style-based generator architectures (e.g. StyleGAN v1, v2) largely promote the controllability and explainability of Generative Adversarial Networks (GANs). Many researchers have applied the pretrained style-based generators to image manipulation and video editing by exploring the correlation between linear interpolation in the latent space and semantic transformation in the synthesized image manifold. However, most previous studies focused on manipulating separate discrete attributes, which is insufficient to animate a still image to generate videos with complex and diverse poses and expressions. In this work, we devise a dual distillation strategy (D2Animator) for generating animated high-resolution face videos conditioned on identities and poses from different images. Specifically, we first introduce a Clustering-based Distiller (CluDistiller) to distill diverse interpolation directions in the latent space, and synthesize identity-consistent faces with various poses and expressions, such as blinking, frowning, looking up/down, etc. Then we propose an Augmentation-based Distiller (AugDistiller) that learns to encode arbitrary face deformation into a combination of interpolation directions via training on augmentation samples synthesized by CluDistiller. Through assembling the two distillation methods, D2Animator can generate high-resolution face animation videos without training on video sequences. Extensive experiments on self-driving, cross-identity and sequence-driving tasks demonstrate the superiority of the proposed D2Animator over existing StyleGAN manipulation and face animation methods in both generation quality and animation fidelity.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1c3a7dc7-ed3e-4fa9-a8ba-28ecf33c4bf5Cited by top-tier papers1
Ask how each one uses itRelated papers
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentQing Chang, Yao-Xiang Ding, Kun ZhouAAAI 2025 · 3 citations
- VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEsMoayed Haji Ali, Andrew Bond, Levent Karacan, Tolga Birdal et al.ICCV 2023 · 3 citations
- StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single VideoLizhen Wang, Xiaochen Zhao, Jingxiang Sun, Yuxiang Zhang et al.SIGGRAPH 2023 · 49 citations
- Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleShuai Tan, Bin Ji, Ye PanAAAI 2024
