AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, Juyong Zhang
摘要
Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing methods that rely on intermediate representations like 2D landmarks or 3D face models to bridge the gap between audio input and video output. Specifically, the feature of input audio signal is directly fed into a conditional implicit function to generate a dynamic neural radiance field, from which a high-fidelity talking-head video corresponding to the audio signal is synthesized using volume rendering. Another advantage of our framework is that not only the head (with hair) region is synthesized as previous methods did, but also the upper body is generated via two individual neural radiance fields. Experimental results demonstrate that our novel framework can (1) produce high-fidelity and natural results, and (2) support free adjustment of audio signals, viewing directions, and background images. Code is available at https://github.com/YudongGuo/AD-NeRF .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper131
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu 等ICCV 2023 · 被引用 192 次
- HeadNeRF: A Realtime NeRF-based Parametric Head ModelYang Hong, Bo Peng, Haiyao Xiao, Ligang Liu 等CVPR 2022 · 被引用 189 次
- Neural Head Avatars from Monocular RGB VideosPhilip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother 等CVPR 2022 · 被引用 173 次
- Generative Neural Articulated Radiance FieldsAlexander W. Bergman, Petr Kellnhofer, Wang Yifan, Eric R. Chan 等NeurIPS 2022 · 被引用 144 次
它引用的顶会 Paper15
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding 等AAAI 2021 · 被引用 88 次
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- D-NeRF: Neural Radiance Fields for Dynamic ScenesAlbert Pumarola, Enric Corona, Gerard Pons-Moll, Francesc Moreno-NoguerCVPR 2021
相关 Paper
- One-Shot High-Fidelity Talking-Head Synthesis with Deformable Neural Radiance FieldWeichuang Li, Longhao Zhang, Dong Wang, Bin Zhao 等CVPR 2023
- Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentRicong Huang, Peiwen Lai, Yipeng Qin, Guanbin LiCVPR 2023
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 等ICLR 2023 · 被引用 33 次
- Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisJiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou 等ICCV 2023 · 被引用 127 次
- Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar ReconstructionGuy Gafni, Justus Thies, Michael Zollhöfer, Matthias NießnerCVPR 2021
