Video-Based Human Pose Regression via Decoupled Space-Time Aggregation
Jijie He, Wenwu Yang
Abstract
By leveraging temporal dependency in video sequences, multi-frame human pose estimation algorithms have demonstrated remarkable results in complicated situations, such as occlusion, motion blur, and video defocus. These algorithms are predominantly based on heatmaps, resulting in high computation and storage requirements per frame, which limits their flexibility and real-time application in video scenarios, particularly on edge devices. In this paper, we develop an efficient and effective video-based human pose regression method, which bypasses intermediate representations such as heatmaps and instead directly maps the input to the output joint coordinates. Despite the inherent spatial correlation among adjacent joints of the human pose, the temporal trajectory of each individual joint exhibits relative independence. In light of this, we propose a novel Decoupled Space-Time Aggregation network (DSTA) to separately capture the spatial contexts between adjacent joints and the temporal cues of each individual joint, thereby avoiding the conflation of spatiotemporal dimensions. Concretely, DSTA learns a dedicated feature token for each joint to facilitate the modeling of their spatiotemporal dependencies. With the proposed joint-wise localawareness attention mechanism, our method is capable of efficiently and flexibly utilizing the spatial dependency of adjacent joints and the temporal dependency of each joint itself. Extensive experiments demonstrate the superiority of our method. Compared to previous regression-based singleframe human pose estimation methods, DSTA significantly enhances performance, achieving an 8.9 mAP improvement on PoseTrack2017. Furthermore, our approach either surpasses or is on par with the state-of-the-art heatmap-based multi-frame human pose estimation methods. Project page: https://github.com/zgspose/DSTA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d40d36b-7ac9-46ba-a5ff-f6fda04a6341Cited by top-tier papers11
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 12 citations
- Causal-Inspired Multitask Learning for Video-Based Human Pose EstimationHaipeng Chen, Sifan Wu, Zhigang Wang, Yifang Yin et al.AAAI 2025 · 7 citations
- SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosYingying Jiao, Zhigang Wang, Sifan Wu, Shaojing Fan et al.AAAI 2025 · 5 citations
- Optimizing Human Pose Estimation Through Focused Human and Joint RegionsYingying Jiao, Zhigang Wang, Zhenguang Liu, Shaojing Fan et al.AAAI 2025 · 4 citations
- High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose EstimationRunyang Feng, Hyung Jin Chang, Tze Ho Elden Tse, Boeun Kim et al.ICCV 2025 · 2 citations
Builds on21
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 360 citations
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
Related papers
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 2 citations
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
