Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis
Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou, Lin Gu
摘要
This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal contribution of spatial regions to guide talking portrait modeling. Specifically, to improve the accuracy of dynamic head reconstruction, a compact and expressive NeRF-based Tri-Plane Hash Representation is introduced by pruning empty spatial regions with three planar hash encoders. For speech audio, we propose a Region Attention Module to generate region-aware condition feature via an attention mechanism. Different from existing methods that utilize an MLP-based encoder to learn the cross-modal relation implicitly, the attention mechanism builds an explicit connection between audio features and spatial regions to capture the priors of local motions. Moreover, a direct and fast Adaptive Pose Encoding is introduced to optimize the head-torso separation problem by mapping the complex transformation of the head pose into spatial coordinates. Extensive experiments demonstrate that our method renders better high-fidelity and audio-lips synchronized talking portrait videos, with realistic details and high efficiency compared to previous methods. Code is available at https://github.com/Fictionarry/ ER-NeRF .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisZiqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 等CVPR 2024 · 被引用 65 次
- GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian SplattingKyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong 等ACM MM 2024 · 被引用 52 次
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen 等ACM MM 2024 · 被引用 28 次
- Learning Dense Correspondence for NeRF-Based Face ReenactmentSonglin Yang, Wei Wang, Yushi Lan, Xiangyu Fan 等AAAI 2024 · 被引用 17 次
- PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head SynthesisYifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo 等AAAI 2025 · 被引用 14 次
它引用的顶会 Paper20
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Plenoxels: Radiance Fields without Neural NetworksSara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen 等CVPR 2022 · 被引用 1,237 次
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano 等CVPR 2022 · 被引用 984 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 被引用 859 次
相关 Paper
- AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisDongze Li, Kang Zhao, Wei Wang, Bo Peng 等AAAI 2024 · 被引用 25 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- Real-Time 3D-Aware Portrait Video RelightingZiqi Cai, Kaiwen Jiang, Shu-Yu Chen, Yu-Kun Lai 等CVPR 2024
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 等ICLR 2023 · 被引用 33 次
- HeadNeRF: A Realtime NeRF-based Parametric Head ModelYang Hong, Bo Peng, Haiyao Xiao, Ligang Liu 等CVPR 2022 · 被引用 189 次
