GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal Modeling
Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu, Chenliang Xu
摘要
Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial interactions between them and result in unnatural and disjointed movements. Additionally, their autoregressive/diffusion-based pipelines show slow generation speed due to dozens of inference steps. To address these two challenges, we propose GestureLSM, a flow-matching-based approach for Co-Speech Gesture Generation with spatial-temporal modeling. Our method i) explicitly model the interaction of tokenized body regions through spatial and temporal attention, for generating coherent full-body gestures. ii) introduce the flow matching to enable more efficient sampling by explicitly modeling the latent velocity space. To overcome the suboptimal performance of flow matching baseline, we propose latent shortcut learning and beta distribution time stamp sampling during training to enhance gesture synthesis quality and accelerate inference. Combining the spatial-temporal modeling and improved flow matching-based framework, GestureLSM achieves state-of-the-art performance on BEAT2 while significantly reducing inference time compared to existing methods, highlighting its potential for enhancing digital humans and embodied agents in real-world applications. Project Page: https://andypinxinliu.github.io/GestureLSM
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 被引用 10 次
- KinMo: Kinematic-Aware Human Motion Understanding and GenerationPengfei Zhang, Pinxin Liu, Pablo Garrido, Hyeongwoo Kim 等ICCV 2025 · 被引用 9 次
- SemGes: Semantics-Aware Co-Speech Gesture Generation Using Semantic Coherence and Relevance LearningLanmiao Liu, Esam Ghaleb, Asli Özyürek, Zerrin YumakICCV 2025 · 被引用 4 次
- LiveGesture: Streamable Co-Speech Gesture Generation ModelMuhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin 等CVPR 2026 · 被引用 4 次
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng 等ICLR 2024 · 被引用 358 次
相关 Paper
- Speech Drives Templates: Co-Speech Gesture Synthesis with Learned TemplatesShenhan Qian, Zhi Tu, Yihao Zhi, Wen Liu 等ICCV 2021 · 被引用 95 次
- DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-Driven Holistic 3D Expression and Gesture GenerationJunming Chen, Yunfei Liu, Jianan Wang, Ailing Zeng 等CVPR 2024
- Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion ModelXu He, Qiaochu Huang, Zhensong Zhang, Zhiwei Lin 等CVPR 2024
- Audio-Driven Co-Speech Gesture Video GenerationXian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 等NeurIPS 2022 · 被引用 77 次
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng 等CVPR 2024 · 被引用 55 次
