Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
Pinxin Liu, Pengfei Zhang, Hyeongwoo Kim, Pablo Garrido, Ari Shapiro, Kyle Olszewski
Abstract
Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the rhythmic or semantic triggers from audio for generating contextualized gesture patterns and achieving pixel-level realism. To address these challenges, we introduce Contextual Gesture, a framework that improves co-speech gesture video generation through three innovative components: (1) a chronological speech-gesture alignment that temporally connects two modalities, (2) a contextualized gesture tokenization that incorporate speech context into motion pattern representation through distillation, and (3) a structure-aware refinement module that employs edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Contextual Gesture not only produces realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications, shown in Fig.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38f04c91-6889-4a25-a7aa-1b7f5ab7d142Cited by top-tier papers4
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- KinMo: Kinematic-Aware Human Motion Understanding and GenerationPengfei Zhang, Pinxin Liu, Pablo Garrido, Hyeongwoo Kim et al.ICCV 2025 · 9 citations
- LiveGesture: Streamable Co-Speech Gesture Generation ModelMuhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin et al.CVPR 2026 · 4 citations
- Bridging Facial Understanding and Animation via Language ModelsLuchuan Song, Pinxin Liu, Haiyang Liu, Zhenchao Jin et al.CVPR 2026
Builds on42
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
Related papers
- Audio-Driven Co-Speech Gesture Video GenerationXian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du et al.NeurIPS 2022 · 77 citations
- Learning Hierarchical Cross-Modal Association for Co-Speech Gesture GenerationXian Liu, Qianyi Wu, Hang Zhou, Yinghao Xu et al.CVPR 2022 · 118 citations
- SemGesture: Synthesizing Semantically Enhanced and Coherent GesturesPengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin XuACM MM 2025 · 2 citations
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin et al.CVPR 2024
- Speech Drives Templates: Co-Speech Gesture Synthesis with Learned TemplatesShenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu et al.ICCV 2021 · 95 citations
