HDG-ODE: A Hierarchical Continuous-Time Model for Human Pose Forecasting
Yucheng Xing, Xin Wang
Abstract
Recently, human pose estimation has attracted more and more attention due to its importance in many real applications. Although many efforts have been put on extracting 2D poses from static images, there are still some severe problems to be solved. A critical one is occlusion, which is more obvious in multi-person scenarios and makes it even more difficult to recover the corresponding 3D poses. When we consider a sequence of images, the temporal correlation among the contexts can be utilized to help us ease the problem, but most of the current works only rely on discrete-time models and estimate the joint locations of all people within a whole sparse graph. In this paper, we propose a new framework, Hierarchical Dynamic Graph Ordinary Differential Equation (HDG-ODE), to tackle the 3D pose forecasting task from 2D skeleton representations in videos. Our framework adopts ODE, a continuous-time model, as the base to predict the 3D joint positions at any time. Considering the structural-property of the skeleton data in representing human poses and the possible irregularity caused by occlusion, we propose the use of dynamic graph convolution as the basic operator. To reduce the computational complexity introduced by the sparsity of the pose graph, our model takes a hierarchical structure where the encoding process at the observation timestamp is done in a cascade manner while the propagation between observations is conducted in parallel. The performance studies on several datasets demonstrate that our model is effective and can out-perform other methods with fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b1ce8c4-75c6-4efa-ba8b-d6c88c227c07Cited by top-tier papers1
Ask how each one uses itBuilds on39
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Optimizing Network Structure for 3D Human Pose EstimationHai Ci, Chunyu Wang, Xiaoxuan Ma, Yizhou WangICCV 2019 · 267 citations
- Single-Stage Multi-Person Pose MachinesXuecheng Nie, Jiashi Feng, Jianfeng Zhang, Shuicheng YanICCV 2019 · 246 citations
- Space-Time-Separable Graph Convolutional Network for Pose ForecastingTheodoros Sofianos, Alessio Sampieri, Luca Franco, Fabio GalassoICCV 2021 · 188 citations
Related papers
- Conditional Directed Graph Convolution for 3D Human Pose EstimationWenbo Hu, Changgong Zhang, Fangneng Zhan, Lei Zhang et al.ACM MM 2021 · 123 citations
- HiPART: Hierarchical Pose AutoRegressive Transformer for Occluded 3D Human Pose EstimationHongwei Zheng, Han Li, Wenrui Dai, Ziyang Zheng et al.CVPR 2025
- Graph Stacked Hourglass Networks for 3D Human Pose EstimationTianhan Xu, Wataru TakanoCVPR 2021
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosYu Cheng, Bo Wang, Bo Yang, Robby T. TanAAAI 2021 · 55 citations
