Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks
Yujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai, Tat-Jen Cham, Junsong Yuan, Nadia Magnenat-Thalmann
Abstract
Despite great progress in 3D pose estimation from single-view images or videos, it remains a challenging task due to the substantial depth ambiguity and severe self-occlusions. Motivated by the effectiveness of incorporating spatial dependencies and temporal consistencies to alleviate these issues, we propose a novel graph-based method to tackle the problem of 3D human body and 3D hand pose estimation from a short sequence of 2D joint detections. Particularly, domain knowledge about the human hand (body) configurations is explicitly incorporated into the graph convolutional operations to meet the specific demand of the 3D pose estimation. Furthermore, we introduce a local-to-global network architecture, which is capable of learning multi-scale features for the graph-based representations. We evaluate the proposed method on challenging benchmark datasets for both 3D hand pose estimation and 3D body pose estimation. Experimental results show that our method achieves state-of-the-art performance on both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers98
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in VideoJinlu Zhang, Zhigang Tu, Jianyu Yang, Yujin Chen et al.CVPR 2022 · 356 citations
- MotionBERT: A Unified Perspective on Learning Human Motion RepresentationsWentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu et al.ICCV 2023 · 322 citations
- Mixup for Node and Graph ClassificationYiwei Wang, Wei Wang, Yuxuan Liang, Yujun Cai et al.WWW 2021 · 220 citations
Builds on2
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageFu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao et al.ICCV 2019 · 178 citations
- SO-HandNet: Self-Organizing Network for 3D Hand Pose Estimation With Semi-Supervised LearningYujin Chen, Zhigang Tu, Liuhao Ge, Dejun Zhang et al.ICCV 2019 · 87 citations
Related papers
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- Sequential Joint Dependency Aware Human Pose Estimation with State Space ModelHanxi Yin, Shaodi You, Jungong Han, Zhixiang ChenAAAI 2025 · 1 citation
- Graph Stacked Hourglass Networks for 3D Human Pose EstimationTianhan Xu, Wataru TakanoCVPR 2021
- Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ICCV 2025 · 1 citation
- HDG-ODE: A Hierarchical Continuous-Time Model for Human Pose ForecastingYucheng Xing, Xin WangICCV 2023 · 5 citations
