Realistic Full-Body Motion Generation from Sparse Tracking with State Space Model
Kun Dong, Jian Xue, Zehai Niu, Xing Lan, Ke Lu, Qingyuan Liu, Xiaoyu Qin
Abstract
In the domain of generative multimedia and interactive experiences, generating realistic and accurate full-body poses from sparse tracking is crucial for many real-world applications, while achieving sequence modeling and efficient motion generation remains challenging. Recently, state space models (SSMs) with efficient hardware-aware designs (i.e., Mamba) have shown great potential for sequence modeling, particularly in temporal contexts. However, processing motion data is still challenging for SSMs. Specifically, the sparsity of input conditions makes motion generation an ill-posed problem. Moreover, the complex structure of the human body further complicates this task. To address these issues, we present Motion Mamba Diffusion (MMD), a novel conditional diffusion model, which effectively utilizes the sequence modeling capability of SSMs and the robust generation ability of diffusion models to track full-body poses accurately. In particular, we design a bidirectional Temporal Mamba Module (TMM) to model motion sequence. Additionally, a Spatial Mamba Module (SMM) is further proposed for feature enhancement within a single frame. Extensive experiments on the large motion capture dataset (AMASS) demonstrate that our proposed approach outperforms the latest methods in terms of accuracy and smoothness, thus providing a crucial advancement for creating realistic virtual avatars in various applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6392a35c-c471-4ae9-a74e-6dc5f583d833Cited by top-tier papers3
- Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationCongyi Fan, Jian Guan, Xuanjia Zhao, Dongli Xu et al.ICCV 2025 · 4 citations
- From Sparse Signal to Smooth Motion: Real-Time Motion Generation with Rolling Prediction ModelsGermán Barquero, Nadine Bertsch, Manojkumar Marramreddy, Carlos Chacón et al.CVPR 2025
- KineST: A Kinematics-guided Spatiotemporal State Space Model for Human Motion Tracking from Sparse SignalsShuting Zhao, Zeyu Xiao, Xinrong ChenAAAI 2026
Related papers
- MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionYusuke Yoshiyasu, Leyuan Sun, Ryusuke SagawaICCV 2025 · 2 citations
- Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion ModelYuming Du, Robin Kips, Albert Pumarola, Sebastian Starke et al.CVPR 2023
- PVMamba: Parallelizing Vision Mamba via Dynamic State AggregationFei Xie, Zhongdao Wang, Weijia Zhang, Chao MaICCV 2025 · 2 citations
- MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space ModelChangcheng Xiao, Qiong Cao, Zhigang Luo, Long LanACM MM 2024 · 31 citations
- DiffusionPose: Markov-Optimized Diffusion Model for Human Pose EstimationZhigang Wang, Zhenguang Liu, Shaojing Fan, Sifan Wu et al.AAAI 2026
