RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
Peng Chen, Xiaobao Wei, Yi Yang, Naiming Yao, Hui Chen, Feng Tian
Abstract
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS)-based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation, with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Our method supports applications such as VR telepresence, social VR, and embodied conversational agents. The socially-aware conditioning can also be extended to other human motion generation tasks. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3794c090-2a9e-4bfd-bf21-6d4f161e6a04Builds on23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisYudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu et al.ICCV 2021 · 510 citations
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang et al.NeurIPS 2024 · 253 citations
Related papers
- GaussianTalker: Real-Time Talking Head Synthesis with 3D Gaussian SplattingKyusun Cho, Joungbin Lee, Heeji Yoon, Yeobin Hong et al.ACM MM 2024 · 52 citations
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsShivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies et al.ICCV 2025 · 8 citations
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingHongyun Yu, Zhan Qu, Qihang Yu, Jianchuan Chen et al.ACM MM 2024 · 28 citations
- SplattingAvatar: Realistic Real-Time Human Avatars With Mesh-Embedded Gaussian SplattingZhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang et al.CVPR 2024 · 92 citations
- Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular VideoSeonghwa Choi, Moonkyeong Choi, Mingyu Jang, Jaekyung Kim et al.ACM MM 2025 · 1 citation
