MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
Zunnan Xu, Yukang Lin, Haonan Han, Sicheng Yang, Ronghui Li, Yachao Zhang, Xiu Li
Abstract
Gesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality. Recent advancements have utilized the diffusion model and attention mechanisms to improve gesture synthesis. However, due to the high computational complexity of these techniques, generating long and diverse sequences with low latency remains a challenge. We explore the potential of state space models (SSMs) to address the challenge, implementing a two-stage modeling strategy with discrete motion priors to enhance the quality of gestures. Leveraging the foundational Mamba block, we introduce MambaTalk, enhancing gesture diversity and rhythm through multimodal integration. Extensive experiments demonstrate that our method matches or exceeds the performance of state-of-the-art models. Our project is publicly available at https://kkakkkka.github.io/MambaTalk
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddf1f221-06e9-43f0-bb8b-55b2430e44fbCited by top-tier papers18
- GestureLSM: Latent Shortcut Based Co-Speech Gesture Generation with Spatial-Temporal ModelingPinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu et al.ICCV 2025 · 54 citations
- MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance GenerationKaixing Yang, Xulong Tang, Ziqiao Peng, Yuxuan Hu et al.NeurIPS 2025 · 25 citations
- MambaTree: Tree Topology is All You Need in State Space ModelYicheng Xiao, Lin Song, Shaoli Huang, Jiangshan Wang et al.NeurIPS 2024 · 21 citations
- MIBURI: Towards Expressive Interactive Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian TheobaltCVPR 2026 · 10 citations
- MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality FusionChencan Fu, Yabiao Wang, Jiangning Zhang, Zhengkai Jiang et al.ACM MM 2024 · 10 citations
Builds on30
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
Related papers
- Realistic Full-Body Motion Generation from Sparse Tracking with State Space ModelKun Dong, Jian Xue, Zehai Niu, Xing Lan et al.ACM MM 2024 · 7 citations
- CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture GenerationFengyi Fang, Sicheng Yang, Wenming YangCVPR 2026 · 4 citations
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic dataTianyi Chen, Pengxiao Lin, Zhiwei Wang, Zhi-Qin John XuNeurIPS 2025 · 4 citations
- DiM-TS: Bridge the Gap Between Selective State Space Models and Time Series for Generative ModelingZihao Yao, Jiankai Zuo, Yaying ZhangAAAI 2026
- Audio-Visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationFa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou et al.ICCV 2025 · 2 citations
