Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Zhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang, Xiaoming Wei, Xunliang Cai, Guanying Chen, Wenhan Luo
Abstract
Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream audio inputs, facing incorrect binding problems between audio and persons. Additionally, they exhibit limitations in instruction-following capabilities. To solve this problem, in this paper, we propose a novel task: Multi-Person Conversational Video Generation, and introduce a new framework, MultiTalk, to address the challenges during multi-person generation. Specifically, for audio injection, we investigate several schemes and propose the Label Rotary Position Embedding (L-RoPE) method to resolve the audio and person binding problem. Furthermore, during training, we observe that partial parameter training and multi-task training are crucial for preserving the instruction-following ability of the base model. MultiTalk achieves superior performance compared to other methods on several datasets, including talking head, talking body, and multi-person datasets, demonstrating the powerful generation capabilities of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Stable Video Infinity: Infinite-Length Video Generation with Error RecyclingWuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao et al.ICLR 2026 · 69 citations
- GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion PriorsXingyilang Yin, Qi Zhang, Jiahao Chang, Ying Feng et al.ICML 2026 · 33 citations
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang et al.ICLR 2026 · 26 citations
- EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationRang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng et al.AAAI 2026 · 24 citations
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio ConditionsZhenzhi Wang, Jiaqi Yang, Jianwen Jiang, Chao Liang et al.ICLR 2026 · 19 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- ReRoPE: Repurposing RoPE for Relative Camera ControlChunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou et al.SIGGRAPH 2026 · 2 citations
- MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingChang Liu, Ye Pan, Chenyang Ding, Susanto Rahardja et al.ACM MM 2025 · 3 citations
- LoL: Longer than Longer, Scaling Video Generation to HourJustin Cui, Jie Wu, Ming Li, Tao Yang et al.CVPR 2026 · 30 citations
- DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationXu Guo, Fulong Ye, Qichao Sun, Liyang Chen et al.ICML 2026 · 16 citations
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationLincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding et al.AAAI 2021 · 88 citations
