Talking Head from Speech Audio using a Pre-trained Image Generator
Mohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. Hogg
Abstract
We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN generator. We model each frame as a point in the latent space of StyleGAN so that a video corresponds to a trajectory through the latent space. Training the network is in two stages. The first stage is to model trajectories in the latent space conditioned on speech utterances. To do this, we use an existing encoder to invert the generator, mapping from each video frame into the latent space. We train a recurrent neural network to map from speech utterances to displacements in the latent space of the image generator. These displacements are relative to the back-projection into the latent space of an identity image chosen from the individuals depicted in the training dataset. In the second stage, we improve the visual quality of the generated videos by tuning the image generator on a single image or a short video of any chosen identity. We evaluate our model on standard measures (PSNR, SSIM, FID and LMD) and show that it significantly outperforms recent state-of-the-art methods on one of two commonly used datasets and gives comparable performance on the other. Finally, we report on ablation experiments that validate the components of the model. The code and videos from experiments can be found at https://mohammedalghamdi.github.io/talking-heads-acm-mm/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c75a15e-8c56-48b6-a1c2-cb6fb293bd6cCited by top-tier papers8
- Talking Head Generation with Probabilistic Audio-to-Visual Diffusion PriorsZhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang et al.ICCV 2023 · 65 citations
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderChenpeng Du, Qi Chen, Tianyu He, Xu Tan et al.ACM MM 2023 · 36 citations
- Say Anything with Any StyleShuai Tan, Bin Ji, Yu Ding, Ye PanAAAI 2024 · 30 citations
- Speech-Driven 3D Face Animation with Composite and Regional Facial MovementsHaozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing et al.ACM MM 2023 · 20 citations
- MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion ModelJin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai et al.ACM MM 2023 · 10 citations
Builds on15
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- GANSpace: Discovering Interpretable GAN ControlsErik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, Sylvain ParisNeurIPS 2020 · 1,049 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
Related papers
- HyperReenact: One-Shot Reenactment via Jointly Learning to Refine and Retarget FacesStella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras et al.ICCV 2023 · 63 citations
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleShuai Tan, Bin Ji, Ye PanAAAI 2024
- Enhancing Identity-Deformation Disentanglement in StyleGAN for One-Shot Face Video Re-EnactmentQing Chang, Yao-Xiang Ding, Kun ZhouAAAI 2025 · 3 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
