SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose Estimation
Wanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao Li
Abstract
Recently, transformer-based methods have been introduced to estimate 3D human pose from multiple views by aggregating the spatial-temporal information of human joints to achieve the lifting of 2D to 3D. However, previous approaches cannot model the inter-frame correspondence of each view's joint individually, nor can they directly consider all view interactions at each time, leading to insufficient learning of multi-view associations. To address this issue, we propose a Spatial-View-Temporal transformer (SVTformer) to decouple spatial-view-temporal information in sequential order for correlation learning and model dependencies between them in a local-to-global manner. SVTformer includes an attended Spatial-View-Temporal (SVT) patch embedding to attentively capture the local features of the input poses and stacked SVT encoders to extract global spatialview-temporal dependencies progressively. Specifically, SVT encoders perform three reconstructions sequentially to attended features with the learning through view decoupling for temporal-enhanced spatial correlation, temporal decoupling for spatial-enhanced view correlation, and another view decoupling for spatial-enhanced temporal relationship. This decoupling-coupling-decoupling multi-view scheme enables us to alternatively model the inter-joint spatial relationships, cross-view dependencies, and temporal motion associations. We evaluate the proposed SVTformer on three popular 3D HPE datasets, and it yields state-of-the-art performance. It effectively deals with ill-posed problems and enhances the accuracy of 3D human pose estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationGeng Chen, Pengfei Ren, Xufeng Jian, Haifeng Sun et al.NeurIPS 2025 · 1 citation
- StructMamPose: From Sequential Perception to Structural Reasoning for 3D Human Pose EstimationJiahong Jiang, Miao Zhang, Jingjing Li, Leiye Liu et al.ICML 2026
Builds on21
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen et al.CVPR 2022 · 356 citations
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo et al.AAAI 2022 · 248 citations
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang et al.ICCV 2019 · 242 citations
Related papers
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose EstimationKangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou et al.ACM MM 2023 · 17 citations
- 3D Human Pose Estimation with Spatio-Temporal Criss-Cross AttentionZhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong et al.CVPR 2023
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Multiple View Geometry Transformers for 3D Human Pose EstimationZiwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu et al.CVPR 2024
