Spatiotemporal-Untrammelled Mixture of Experts for Multi-Person Motion Prediction
Zheng Yin, Chengjian Li, Xiangbo Shu, Meiqi Cao, Rui Yan, Jinhui Tang
Abstract
Comprehensively and flexibly capturing the complex spatiotemporal dependencies of human motion is critical for multiperson motion prediction. Existing methods grapple with two primary limitations: i) Inflexible spatiotemporal representation due to reliance on positional encodings for capturing spatiotemporal information. ii) High computational costs stemming from the quadratic time complexity of conventional attention mechanisms. To overcome these limitations, we propose the Spatiotemporal-Untrammelled Mixture of Experts (ST-MoE), which flexibly explores complex spatio-temporal dependencies in human motion and significantly reduces computational cost. To adaptively mine complex spatio-temporal patterns from human motion, our model incorporates four distinct types of spatiotemporal experts, each specializing in capturing different spatial or temporal dependencies. To reduce the potential computational overhead while integrating multiple experts, we introduce bidirectional spatiotemporal Mamba as experts, each sharing bidirectional temporal and spatial Mamba in distinct combinations to achieve model efficiency and parameter economy. Extensive experiments on four multi-person benchmark datasets demonstrate that our approach not only outperforms state-of-art in accuracy but also reduces model parameter by 41.38% and achieves a 3.6× speedup in training. The code is available at https://github.com/alanyz106/ST-MoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffd746f8-e940-4fa0-89c5-e794539c953bBuilds on24
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionLingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 252 citations
Related papers
- STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series PredictionHaolong Chen, Liang Zhang, Zhengyuan Xin, Guangxu ZhuKDD 2026 · 1 citation
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 15 citations
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationRunyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim et al.ICCV 2025 · 2 citations
- Efficient Multi-Person Motion Prediction by Lightweight Spatial and Temporal InteractionsYuanhong Zheng, Ruixuan Yu, Jian SunICCV 2025
- Trajectory Mamba: Efficient Attention-Mamba Forecasting Model Based on Selective SSMYizhou Huang, Yihua Cheng, Kezhi WangCVPR 2025
