GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
Hongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen, Zhonghua Jiang, Zhiwen Chen, Shengyu Zhang, Jimin Xu, Fei Wu, Chengfei Lv, Gang Yu
摘要
Recent works on audio-driven talking head synthesis using Neural Radiance Fields (NeRF) have achieved impressive results. However, due to inadequate pose and expression control caused by NeRF implicit representation, these methods still have some limitations, such as unsynchronized or unnatural lip movements, and visual jitter and artifacts. In this paper, we propose GaussianTalker, a novel method for audio-driven talking head synthesis based on 3D Gaussian Splatting. With the explicit representation property of 3D Gaussians, intuitive control of the facial motion is achieved by binding Gaussians to 3D facial models. GaussianTalker consists of two modules, Speaker-specific Motion Translator and Dynamic Gaussian Renderer. Speaker-specific Motion Translator achieves accurate lip movements specific to the target speaker through universalized audio feature extraction and customized lip motion generation. Dynamic Gaussian Renderer introduces Speaker-specific BlendShapes to enhance facial detail representation via a latent pose, delivering stable and realistic rendered videos. Extensive experimental results suggest that GaussianTalker outperforms existing state-of-the-art methods in talking head synthesis, delivering precise lip synchronization and exceptional visual quality. Our method achieves rendering speeds of 130 FPS on NVIDIA RTX4090 GPU, significantly exceeding the threshold for real-time rendering performance, and can potentially be deployed on other hardware platforms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache OptimizationZhonghua Jiang, Kui Chen, Kunxi Li, Keting Yin 等AAAI 2026 · 被引用 4 次
- MixedGaussianAvatar: Realistically and Geometrically Accurate Head Avatar via Mixed 2D-3D GaussiansPeng Chen, Xiaobao Wei, Qingpo Wuwu, Xinyi Wang 等ACM MM 2025 · 被引用 2 次
- DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsXiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang 等ICCV 2025 · 被引用 2 次
- DashGaussian: Optimizing 3D Gaussian Splatting in 200 SecondsYouyu Chen, Junjun Jiang, Kui Jiang, Xiao Tang 等CVPR 2025
- InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoJiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng 等CVPR 2025
它引用的顶会 Paper15
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz 等ICCV 2021 · 被引用 1,442 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
- EMOCA: Emotion Driven Monocular Face Capture and AnimationRadek Danecek, Michael J. Black, Timo BolkartCVPR 2022 · 被引用 180 次
相关 Paper
- GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian SplattingKyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong 等ACM MM 2024 · 被引用 52 次
- GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian SplattingXiaobao Wei, Peng Chen, Guangyu Li, Ming Lu 等ICCV 2025 · 被引用 2 次
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsShivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies 等ICCV 2025 · 被引用 8 次
- GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh LearningYe Yuan, Xueting Li, Yangyi Huang, Shalini De Mello 等CVPR 2024
- Hierarchically Controlled Deformable 3D Gaussians for Talking Head SynthesisZhenhua Wu, Linxuan Jiang, Xiang Li, Chaowei Fang 等AAAI 2025 · 被引用 2 次
