AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head Synthesis
Dongze Li, Kang Zhao, Wei Wang, Bo Peng, Yingya Zhang, Jing Dong, Tieniu Tan
摘要
Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerShurong Yang, Huadong Li, Juhao Wu, Minhao Jing 等AAAI 2025 · 被引用 12 次
- InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoJiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng 等CVPR 2025
它引用的顶会 Paper15
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang 等ICCV 2021 · 被引用 1,024 次
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano 等CVPR 2022 · 被引用 984 次
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 被引用 687 次
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 被引用 622 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
相关 Paper
- Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisJiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou 等ICCV 2023 · 被引用 127 次
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi 等CVPR 2026 · 被引用 5 次
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 等ICLR 2023 · 被引用 33 次
- Context-Aware Talking-Head Video EditingSonglin Yang, Wei Wang, Jun Ling, Bo Peng 等ACM MM 2023 · 被引用 10 次
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
