SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose Estimation
Wanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao Li
摘要
Recently, transformer-based methods have been introduced to estimate 3D human pose from multiple views by aggregating the spatial-temporal information of human joints to achieve the lifting of 2D to 3D. However, previous approaches cannot model the inter-frame correspondence of each view's joint individually, nor can they directly consider all view interactions at each time, leading to insufficient learning of multi-view associations. To address this issue, we propose a Spatial-View-Temporal transformer (SVTformer) to decouple spatial-view-temporal information in sequential order for correlation learning and model dependencies between them in a local-to-global manner. SVTformer includes an attended Spatial-View-Temporal (SVT) patch embedding to attentively capture the local features of the input poses and stacked SVT encoders to extract global spatialview-temporal dependencies progressively. Specifically, SVT encoders perform three reconstructions sequentially to attended features with the learning through view decoupling for temporal-enhanced spatial correlation, temporal decoupling for spatial-enhanced view correlation, and another view decoupling for spatial-enhanced temporal relationship. This decoupling-coupling-decoupling multi-view scheme enables us to alternatively model the inter-joint spatial relationships, cross-view dependencies, and temporal motion associations. We evaluate the proposed SVTformer on three popular 3D HPE datasets, and it yields state-of-the-art performance. It effectively deals with ill-posed problems and enhances the accuracy of 3D human pose estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationGeng Chen, Pengfei Ren, Xufeng Jian, Haifeng Sun 等NeurIPS 2025 · 被引用 1 次
- StructMamPose: From Sequential Perception to Structural Reasoning for 3D Human Pose EstimationJiahong Jiang, Miao Zhang, Jingjing Li, Leiye Liu 等ICML 2026
它引用的顶会 Paper21
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 被引用 419 次
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen 等CVPR 2022 · 被引用 356 次
- Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on TransformerTao Wang, Hong Liu, Pinhao Song, Tianyu Guo 等AAAI 2022 · 被引用 248 次
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang 等ICCV 2019 · 被引用 242 次
相关 Paper
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou 等AAAI 2024 · 被引用 14 次
- Efficient Hierarchical Multi-view Fusion Transformer for 3D Human Pose EstimationKangkang Zhou, Lijun Zhang, Feng Lu, Xiang-Dong Zhou 等ACM MM 2023 · 被引用 17 次
- 3D Human Pose Estimation with Spatio-Temporal Criss-Cross AttentionZhenhua Tang, Zhaofan Qiu, Yanbin Hao, Richang Hong 等CVPR 2023
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- Multiple View Geometry Transformers for 3D Human Pose EstimationZiwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu 等CVPR 2024
