Talking Together: Synthesizing Co-Located 3D Conversations from Audio
Mengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu, Yinda Zhang, Luchuan Song, Rohit Pandey, Sean Fanello, Zeng Huang
Abstract
We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied ``talking heads'' akin to a video conference call, our work is the first to explicitly model the dynamic 3D spatial relationship—including relative position, orientation, and mutual gaze—that is crucial for realistic in-person dialogues. Our system synthesizes the full performance of both individuals, including precise lip-sync, and uniquely allows their relative head poses to be controlled via textual descriptions. To achieve this, we propose a dual-stream architecture where each stream is responsible for one participant's output. We employ speaker's role embeddings and inter-speaker cross-attention mechanisms are designed to disentangle the mixed audio and model the interaction. Furthermore, we introduce a novel eye gaze loss to promote natural, mutual eye contact. To power our data-hungry approach, we introduce a novel pipeline to curate a large-scale conversational dataset consisting of over 2 million dyadic pairs from in-the-wild videos. Our method generates fluid, controllable, and spatially aware dyadic animations suitable for immersive applications in VR and telepresence, significantly outperforming existing baselines in perceived realism and interaction coherence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7892a1bf-3f7a-4071-9a2e-c4a23b20db6bBuilds on28
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- FaceFormer: Speech-Driven 3D Facial Animation with TransformersYingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang et al.CVPR 2022 · 218 citations
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu et al.ICCV 2023 · 192 citations
- Imitator: Personalized Speech-driven 3D Facial AnimationBalamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker et al.ICCV 2023 · 98 citations
Related papers
- The eyes have it: an integrated eye and face model for photorealistic facial animationGabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi et al.SIGGRAPH 2020 · 54 citations
- Learning to Listen: Modeling Non-Deterministic Dyadic Facial MotionEvonne Ng, Hanbyul Joo, Liwen Hu, Hao Li et al.CVPR 2022 · 87 citations
- DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsZiqiao Peng, Yanbo Fan, Haoyu Wu, Xuan Wang et al.CVPR 2025
- RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationPeng Chen, Xiaobao Wei, Yi Yang, Naiming Yao et al.IEEE VR 2026
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou et al.CVPR 2022 · 114 citations
