Audio-Driven Co-Speech Gesture Video Generation
Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du, Wayne Wu, Dahua Lin, Ziwei Liu
Abstract
Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the image domain remains unsolved. In this work, we formally define and study this challenging problem of audio-driven co-speech gesture video generation, i.e., using a unified framework to generate speaker image sequence driven by speech audio. Our key insight is that the co-speech gestures can be decomposed into common motion patterns and subtle rhythmic dynamics. To this end, we propose a novel framework, Audio-driveN Gesture vIdeo gEneration (ANGIE), to effectively capture the reusable co-speech gesture patterns as well as fine-grained rhythmic movements. To achieve high-fidelity image sequence generation, we leverage an unsupervised motion representation instead of a structural human body prior (e.g., 2D skeletons). Specifically, 1) we propose a vector quantized motion extractor (VQ-Motion Extractor) to summarize common co-speech gesture patterns from implicit motion representation to codebooks. 2) Moreover, a co-speech gesture GPT with motion refinement (Co-Speech GPT) is devised to complement the subtle prosodic motion details. Extensive experiments demonstrate that our framework renders realistic and vivid co-speech gesture video. Demo video and more resources can be found in: https://alvinliu0.github.io/projects/ANGIE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3957014e-d9fa-4d0f-a684-267a3b35ceb0Cited by top-tier papers22
- HyperHuman: Hyper-Realistic Human Generation with Latent Structural DiffusionXian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov et al.ICLR 2024 · 81 citations
- GAIA: Zero-shot Talking Avatar GenerationTianyu He, Junliang Guo, Runyi Yu, Yuchi Wang et al.ICLR 2024 · 51 citations
- LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture GenerationYihao Zhi, Xiaodong Cun, Xuelin Chen, Xi Shen et al.ICCV 2023 · 48 citations
- HumanGaussian: Text-Driven 3D Human Generation with Gaussian SplattingXian Liu, Xiaohang Zhan, Jiaxiang Tang, Ying Shan et al.CVPR 2024 · 42 citations
- Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationLuchuan Song, Guojun Yin, Zhenchao Jin, Xiaoyi Dong et al.ICCV 2023 · 19 citations
Builds on11
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementAlexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre et al.ICCV 2021 · 272 citations
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic MemoryLi Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin et al.CVPR 2022 · 170 citations
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
Related papers
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
- Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture RepresentationPinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Garrido et al.ACM MM 2025
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture GenerationXian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu et al.CVPR 2022 · 118 citations
- Co-Speech Gesture Video Generation with Implicit Motion-Audio EntanglementXinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu et al.CVPR 2025
- Democratizing High-Fidelity Co-Speech Gesture Video GenerationXu Yang, Shaoli Huang, Shenbo Xie, Xuelin Chen et al.ICCV 2025 · 1 citation
