AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, Juyong Zhang
Abstract
Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing methods that rely on intermediate representations like 2D landmarks or 3D face models to bridge the gap between audio input and video output. Specifically, the feature of input audio signal is directly fed into a conditional implicit function to generate a dynamic neural radiance field, from which a high-fidelity talking-head video corresponding to the audio signal is synthesized using volume rendering. Another advantage of our framework is that not only the head (with hair) region is synthesized as previous methods did, but also the upper body is generated via two individual neural radiance fields. Experimental results demonstrate that our novel framework can (1) produce high-fidelity and natural results, and (2) support free adjustment of audio signals, viewing directions, and background images. Code is available at https://github.com/YudongGuo/AD-NeRF .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers131
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- HeadNeRF: A Realtime NeRF-based Parametric Head ModelYang Hong, Bo Peng, Haiyao Xiao, Ligang Liu et al.CVPR 2022 · 189 citations
- Neural Head Avatars from Monocular RGB VideosPhilip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother et al.CVPR 2022 · 173 citations
- Generative Neural Articulated Radiance FieldsAlexander W. Bergman, Petr Kellnhofer, Wang Yifan, Eric R. Chan et al.NeurIPS 2022 · 144 citations
Builds on15
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- D-NeRF: Neural Radiance Fields for Dynamic ScenesAlbert Pumarola, Enric Corona, Gerard Pons-Moll, Francesc Moreno-NoguerCVPR 2021
Related papers
- One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldWeichuang Li, Longhao Zhang, Dong Wang, Bin Zhao et al.CVPR 2023
- Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentRicong Huang, Peiwen Lai, Yipeng Qin, Guanbin LiCVPR 2023
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu et al.ICLR 2023 · 33 citations
- Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisJiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou et al.ICCV 2023 · 127 citations
- Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar ReconstructionGuy Gafni, Justus Thies, Michael Zollhöfer, Matthias NießnerCVPR 2021
