GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian Splatting
Kyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong, Jaehoon Ko, Sangjun Ahn, Seungryong Kim
Abstract
This paper proposes GaussianTalker, a novel framework for real-time generation of pose-controllable talking heads. It leverages the fast rendering capabilities of 3D Gaussian Splatting (3DGS) while addressing the challenges of directly controlling 3DGS with speech audio. GaussianTalker constructs a single 3DGS representation of the head and deforms it in sync with the audio. A key insight is to encode the 3D Gaussian attributes into a shared implicit feature representation, where it is merged with audio features to manipulate each Gaussian attribute. This design exploits the spatial information of the head and enforces interactions between neighboring points. The feature embeddings are then fed to a spatial-audio attention module, which predicts frame-wise offsets for the attributes of each Gaussian. This method is more stable than previous concatenation or multiplication approaches for manipulating the numerous Gaussians and their intricate parameters. Overall, GaussianTalker offers a promising approach for real-time generation of high-quality pose-controllable talking heads. pecifically, GaussianTalker achieves a remarkable rendering speed up to 120 FPS, surpassing previous benchmarks. Our demo video and code can be found at https://ku-cvlab.github.io/GaussianTalker/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd47729e-82b5-4be1-bc9c-d0010c3615abCited by top-tier papers10
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsShivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies et al.ICCV 2025 · 8 citations
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi et al.CVPR 2026 · 5 citations
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer et al.ICCV 2025 · 3 citations
- Hierarchically Controlled Deformable 3D Gaussians for Talking Head SynthesisZhenhua Wu, Linxuan Jiang, Xiang Li, Chaowei Fang et al.AAAI 2025 · 2 citations
- GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarSeungjun Moon, Hah Min Lew, Seungeun Lee, Ji-Su Kang et al.ICCV 2025 · 2 citations
Builds on25
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian SplattingZeyu Yang, Hongye Yang, Zijie Pan, Li ZhangICLR 2024 · 529 citations
Related papers
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen et al.ACM MM 2024 · 28 citations
- RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationPeng Chen, Xiaobao Wei, Yi Yang, Naiming Yao et al.IEEE VR 2026
- GauHuman: Articulated Gaussian Splatting from Monocular Human VideosShoukang Hu, Tao Hu, Ziwei LiuCVPR 2024
- GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh LearningYe Yuan, Xueting Li, Yangyi Huang, Shalini De Mello et al.CVPR 2024
- DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsXiaoxi Liang, Yanbo Fan, Qiya Yang, Xuan Wang et al.ICCV 2025 · 2 citations
