HumMUSS: Human Motion Understanding Using State Space Models
Arnab Kumar Mondal, Stefano Alletto, Denis Tomè
Abstract
Understanding human motion from video is essential for a range of applications, including pose estimation, mesh recovery and action recognition. While state-of-the-art methods predominantly rely on transformer-based architectures, these approaches have limitations in practical scenarios. Transformers are slower when sequentially predicting on a continuous stream of frames in real-time, and do not generalize to new frame rates. In light of these constraints, we propose a novel attention-free spatiotemporal model for human motion understanding building upon recent advancements in state space models. Our model not only matches the performance of transformer-based models in various motion understanding tasks but also brings added benefits like adaptability to different video frame rates and enhanced training speed when working with longer sequences of keypoints. Moreover, the proposed model supports both offline and real-time applications. For real-time sequential prediction, our model is both memory efficient and several times faster than transformer-based approaches while maintaining their high accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN NetworkXinyi Zhang, Qiqi Bao, Qinpeng Cui, Wenming Yang et al.AAAI 2025 · 18 citations
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 15 citations
- MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentYachun Mi, Yu Li, Weicheng Meng, Chaofeng Chen et al.ICCV 2025 · 1 citation
- UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and EditingYiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan et al.CVPR 2025
Builds on49
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
Related papers
- MixSynthFormer: A Transformer Encoder-like Structure with Mixed Synthetic Self-attention for Efficient Human Pose EstimationYuran Sun, Alan William Dougherty, Zhuoying Zhang, Yi-King Choi et al.ICCV 2023 · 6 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- 15 Keypoints Is All You NeedMichael Snower, Asim Kadav, Farley Lai, Hans Peter GrafCVPR 2020
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen et al.CVPR 2022 · 356 citations
