3D Human Pose Estimation with Spatial and Temporal Transformers
Ce Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang, Chen Chen, Zhengming Ding
Abstract
Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation. However, in the field of human pose estimation, convolutional architectures still remain dominant. In this work, we present PoseFormer, a purely transformer-based approach for 3D human pose estimation in videos without convolutional architectures involved. Inspired by recent developments in vision transformers, we design a spatial-temporal transformer structure to comprehensively model the human joint relations within each frame as well as the temporal correlations across frames, then output an accurate 3D human pose of the center frame. We quantitatively and qualitatively evaluate our method on two popular and standard benchmark datasets: Human3.6M and MPI-INF-3DHP. Extensive experiments show that PoseFormer achieves state-of-the-art performance on both datasets. Code is available at https://github.com/zczcwh/PoseFormer
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c499872f-7b3d-4c4f-9821-a2bb0e280251Cited by top-tier papers103
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen et al.CVPR 2022 · 356 citations
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingYuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan et al.NeurIPS 2022 · 222 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 267 citations
Related papers
- 3D Human Pose Estimation with Spatio-Temporal Criss-Cross AttentionZhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong et al.CVPR 2023
- PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Pichao Wang et al.CVPR 2023
- MixSynthFormer: A Transformer Encoder-like Structure with Mixed Synthetic Self-attention for Efficient Human Pose EstimationYuran Sun, Alan William Dougherty, Zhuoying Zhang, Yi-King Choi et al.ICCV 2023 · 6 citations
- KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose EstimationJihua Peng, Yanghong Zhou, P. Y. MokCVPR 2024 · 67 citations
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 15 citations
