Learning Dynamics via Graph Neural Networks for Human Pose Estimation and Tracking
Yiding Yang, Zhou Ren, Haoxiang Li, Chunluan Zhou, Xinchao Wang, Gang Hua
Abstract
Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably prone to missed detections especially in heavilycluttered scenes, since this tracking-by-detection paradigm is, by nature, largely dependent on visual evidences that are absent in the case of occlusion. In this paper, we propose a novel online approach to learning the pose dynamics, which are independent of pose detections in current fame, and hence may serve as a robust estimation even in challenging scenarios including occlusion. Specifically, we derive this prediction of dynamics through a graph neural network (GNN) that explicitly accounts for both spatialtemporal and visual information. It takes as input the historical pose tracklets and directly predicts the corresponding poses in the following frame for each tracklet. The predicted poses will then be aggregated with the detected poses, if any, at the same frame so as to produce the final pose, potentially recovering the occluded joints missed by the estimator. Experiments on PoseTrack 2017 and Pose-Track 2018 datasets demonstrate that the proposed method achieves results superior to the state of the art on both human pose estimation and tracking tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu et al.CVPR 2022 · 76 citations
- PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose TrackingAndreas Doering, Di Chen, Shanshan Zhang, Bernt Schiele et al.CVPR 2022 · 47 citations
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.ICCV 2021 · 46 citations
- Stochastic Partial Swap: Enhanced Model Generalization and Interpretability for Fine-grained RecognitionShaoli Huang, Xinchao Wang, Dacheng TaoICCV 2021 · 46 citations
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
Builds on9
- Factorizable Graph Convolutional NetworksYiding Yang, Zunlei Feng, Mingli Song, Xinchao WangNeurIPS 2020 · 175 citations
- DGCN: Dynamic Graph Convolutional Network for Efficient Multi-Person Pose EstimationZhongwei Qiu, Kai Qiu, Jianlong Fu, Dongmei FuAAAI 2020 · 52 citations
- 15 Keypoints Is All You NeedMichael Snower, Asim Kadav, Farley Lai, Hans Peter GrafCVPR 2020
- Distilling Knowledge From Graph Convolutional NetworksYiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao et al.CVPR 2020
- Dynamic Graph Message Passing NetworksLi Zhang, Dan Xu, Anurag Arnab, Philip H. S. TorrCVPR 2020
Related papers
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosYu Cheng, Bo Wang, Bo Yang, Robby T. TanAAAI 2021 · 55 citations
- GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object TrackerYaxuan Hu, Jie Hua, Gang Wu, Yuhong Yang et al.AAAI 2026
- CoMotion: Concurrent Multi-person 3D MotionAlejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter et al.ICLR 2025
