GaussianSpeech: Audio-Driven Personalized 3D Gaussian Avatars
Shivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies, Angela Dai, Matthias Nießner
摘要
We introduce GaussianSpeech 1 , a novel approach that synthesizes high-fidelity animation sequences of photorealistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including skin furrowing and finer-scale facial movements, we propose to couple speech signal with 3D Gaussian splatting to create realistic, temporally coherent motion sequences. We propose a compact and efficient 3DGS-based avatar representation that generates expression-dependent color and leverages wrinkle-and perceptually-based losses to synthesize facial details. To enable sequence modeling of 3D Gaussian splats with audio, we devise an audio-conditioned transformer model capable of extracting lip and expression features directly from audio input. Due to the absence of high-quality dataset of talking humans in correspondence with audio, we captured a new large-scale multi-view dataset of audio-visual sequences of talking humans with native English accents and diverse facial geometry. GaussianSpeech consistently achieves state-of-the-art quality with visually natural motion, while encompassing diverse facial expressions and styles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadChang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou 等CVPR 2026 · 被引用 1 次
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view GenerationAviral Chharia, Fernando De la TorreCVPR 2026
它引用的顶会 Paper29
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu 等ICCV 2021 · 被引用 510 次
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre 等ICCV 2021 · 被引用 272 次
相关 Paper
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen 等ACM MM 2024 · 被引用 28 次
- ScaffoldAvatar: High-Fidelity Gaussian Avatars with Patch ExpressionsShivangi Aneja, Sebastian Weiss, Irene Baeza, Prashanth Chandran 等SIGGRAPH 2025 · 被引用 6 次
- 3D Gaussian Blendshapes for Head Avatar AnimationShengjie Ma, Yanlin Weng, Tianjia Shao, Kun ZhouSIGGRAPH 2024 · 被引用 57 次
- GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian SplattingKyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong 等ACM MM 2024 · 被引用 52 次
- RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationPeng Chen, Xiaobao Wei, Yi Yang, Naiming Yao 等IEEE VR 2026
