AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head Synthesis
Dongze Li, Kang Zhao, Wei Wang, Bo Peng, Yingya Zhang, Jing Dong, Tieniu Tan
Abstract
Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies. However, when it comes to few-shot talking head generation, a practical scenario where only few seconds of talking video is available for one identity, two limitations emerge: 1) they either have no base model, which serves as a facial prior for fast convergence, or ignore the importance of audio when building the prior; 2) most of them overlook the degree of correlation between different face regions and audio, e.g., mouth is audio related, while ear is audio independent. In this paper, we present Audio Enhanced Neural Radiance Field (AE-NeRF) to tackle the above issues, which can generate realistic portraits of a new speaker with few-shot dataset. Specifically, we introduce an Audio Aware Aggregation module into the feature fusion stage of the reference scheme, where the weight is determined by the similarity of audio between reference and target image. Then, an Audio-Aligned Face Generation strategy is proposed to model the audio related and audio independent regions respectively, with a dual-NeRF framework. Extensive experiments have shown AE-NeRF surpasses the state-of-the-art on image fidelity, audio-lip synchronization, and generalization ability, even in limited training set or training iterations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a64e4da-c5eb-4267-9b36-cd719ca63730Cited by top-tier papers2
- MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion TransformerShurong Yang, Huadong Li, Juhao Wu, Minhao Jing et al.AAAI 2025 · 12 citations
- InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoJiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng et al.CVPR 2025
Builds on15
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- Few-Shot Adversarial Learning of Realistic Neural Talking Head ModelsEgor Zakharov, Aliaksandra Shysheya, Egor Burkov, Victor S. LempitskyICCV 2019 · 687 citations
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 622 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
Related papers
- Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisJiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou et al.ICCV 2023 · 127 citations
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi et al.CVPR 2026 · 5 citations
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu et al.ICLR 2023 · 33 citations
- Context-Aware Talking-Head Video EditingSonglin Yang, Wei Wang, Jun Ling, Bo Peng et al.ACM MM 2023 · 10 citations
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu et al.CVPR 2024 · 65 citations
