GaussianSpeech: Audio-Driven Personalized 3D Gaussian Avatars
Shivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies, Angela Dai, Matthias Nießner
Abstract
We introduce GaussianSpeech 1 , a novel approach that synthesizes high-fidelity animation sequences of photorealistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including skin furrowing and finer-scale facial movements, we propose to couple speech signal with 3D Gaussian splatting to create realistic, temporally coherent motion sequences. We propose a compact and efficient 3DGS-based avatar representation that generates expression-dependent color and leverages wrinkle-and perceptually-based losses to synthesize facial details. To enable sequence modeling of 3D Gaussian splats with audio, we devise an audio-conditioned transformer model capable of extracting lip and expression features directly from audio input. Due to the absence of high-quality dataset of talking humans in correspondence with audio, we captured a new large-scale multi-view dataset of audio-visual sequences of talking humans with native English accents and diverse facial geometry. GaussianSpeech consistently achieves state-of-the-art quality with visually natural motion, while encompassing diverse facial expressions and styles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e2b1909-9c0d-4fd8-98cc-e67e76635267Cited by top-tier papers2
- EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadChang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou et al.CVPR 2026 · 1 citation
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view GenerationAviral Chharia, Fernando De la TorreCVPR 2026
Builds on29
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
Related papers
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen et al.ACM MM 2024 · 28 citations
- ScaffoldAvatar: High-Fidelity Gaussian Avatars with Patch ExpressionsShivangi Aneja, Sebastian Weiss, Irene Baeza, Prashanth Chandran et al.SIGGRAPH 2025 · 6 citations
- 3D Gaussian Blendshapes for Head Avatar AnimationShengjie Ma, Yanlin Weng, Tianjia Shao, Kun ZhouSIGGRAPH 2024 · 57 citations
- GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian SplattingKyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong et al.ACM MM 2024 · 52 citations
- RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationPeng Chen, Xiaobao Wei, Yi Yang, Naiming Yao et al.IEEE VR 2026
