One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing
Ting-Chun Wang, Arun Mallya, Ming-Yu Liu
Abstract
NVIDIA Corporation (a) Original video (b) Compressed videos at the same bit-rate (c) Our re-rendered novel-view results Figure 1: Our method can re-create a talking-head video using only a single source image (e.g., the first frame) and a sequence of unsupervisedly-learned 3D keypoints, representing motions in the video. Our novel keypoint representation provides a compact representation of the video that is 10ˆmore efficient than the H.264 baseline can provide. A novel 3D keypoint decomposition scheme allows re-rendering the talking-head video under different poses, simulating often missed face-to-face video conferencing experiences. Video versions of the paper figures and additional results are available at our project page, https://nvlabs.github.io/face-vid2vid .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers179
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingYurui Ren, Ge Li, Yuanqi Chen, Thomas H. Li et al.ICCV 2021 · 284 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- Latent Image Animator: Learning to Animate Images via Latent Space NavigationYaohui Wang, Di Yang, François Brémond, Antitza DantchevaICLR 2022 · 219 citations
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 196 citations
Builds on20
- Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?Rameen Abdal, Yipeng Qin, Peter WonkaICCV 2019 · 1,195 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Few-Shot Unsupervised Image-to-Image TranslationMing-Yu Liu, Xun Huang, Arun Mallya, Tero Karras et al.ICCV 2019 · 668 citations
Related papers
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben et al.CVPR 2020
- IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCYuan Li, Ziqian Bai, Feitong Tan, Zhaopeng Cui et al.CVPR 2025
- Occlusion-Insensitive Talking Head Video Generation via Facelet CompensationYuhui Deng, Yuqin Lu, Yangyang Xu, Yongwei Nie et al.AAAI 2025 · 3 citations
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationWenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang et al.CVPR 2023
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 9 citations
