PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation
Qitao Zhao, Ce Zheng, Mengyuan Liu, Pichao Wang, Chen Chen
摘要
Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics across frames with cascaded transformer layers and has achieved impressive performance. However, in real scenarios, the performance of PoseFormer and its follow-ups is limited by two factors: (a) The length of the input joint sequence; (b) The quality of 2D joint detection. Existing methods typically apply self-attention to all frames of the input sequence, causing a huge computational burden when the frame number is increased to obtain advanced estimation accuracy, and they are not robust to noise naturally brought by the limited capability of 2D joint detectors. In this paper, we propose PoseFormerV2, which exploits a compact representation of lengthy skeleton sequences in the frequency domain to efficiently scale up the receptive field and boost robustness to noisy 2D joint detection. With minimum modifications to PoseFormer, the proposed method effectively fuses features both in the time domain and frequency domain, enjoying a better speed-accuracy trade-off than its precursor. Extensive experiments on two benchmark datasets (i.e., Human3.6M and MPI-INF-3DHP) demonstrate that the proposed approach significantly outperforms the original PoseFormer and other transformer-based variants. Code is released at https://github.com/ QitaoZhao/PoseFormerV2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose EstimationJihua Peng, Yanghong Zhou, P. Y. MokCVPR 2024 · 被引用 67 次
- A Single 2D Pose with Context is Worth Hundreds for 3D Human Pose EstimationQitao Zhao, Ce Zheng, Mengyuan Liu, Chen ChenNeurIPS 2023 · 被引用 40 次
- Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal DenoiserQingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 等AAAI 2024 · 被引用 39 次
- FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion ModelsJinglin Xu, Yijie Guo, Yuxin PengCVPR 2024 · 被引用 39 次
- A Dual-Augmentor Framework for Domain Generalization in 3D Human Pose EstimationQucheng Peng, Ce Zheng, Chen ChenCVPR 2024 · 被引用 38 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- ExtPose: Robust and Coherent Pose Estimation by Extending ViTsRongyu Chen, Li'an Zhuo, Linlin Yang, Qi Wang 等ICML 2025
- HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose EstimationHongwei Zheng, Han Li, Wenrui Dai, Ziyang Zheng 等CVPR 2025
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen 等CVPR 2022 · 被引用 356 次
- SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose EstimationWanruo Zhang, Mengyuan Liu, Hong Liu, Wenhao LiAAAI 2025 · 被引用 4 次
