Learning Dynamics via Graph Neural Networks for Human Pose Estimation and Tracking
Yiding Yang, Zhou Ren, Haoxiang Li, Chunluan Zhou, Xinchao Wang, Gang Hua
摘要
Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably prone to missed detections especially in heavilycluttered scenes, since this tracking-by-detection paradigm is, by nature, largely dependent on visual evidences that are absent in the case of occlusion. In this paper, we propose a novel online approach to learning the pose dynamics, which are independent of pose detections in current fame, and hence may serve as a robust estimation even in challenging scenarios including occlusion. Specifically, we derive this prediction of dynamics through a graph neural network (GNN) that explicitly accounts for both spatialtemporal and visual information. It takes as input the historical pose tracklets and directly predicts the corresponding poses in the following frame for each tracklet. The predicted poses will then be aggregated with the detected poses, if any, at the same frame so as to produce the final pose, potentially recovering the occluded joints missed by the estimator. Experiments on PoseTrack 2017 and Pose-Track 2018 datasets demonstrate that the proposed method achieves results superior to the state of the art on both human pose estimation and tracking tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 等CVPR 2022 · 被引用 76 次
- PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose TrackingAndreas Doering, Di Chen, Shanshan Zhang, Bernt Schiele 等CVPR 2022 · 被引用 47 次
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song 等ICCV 2021 · 被引用 46 次
- Stochastic Partial Swap: Enhanced Model Generalization and Interpretability for Fine-grained RecognitionShaoli Huang, Xinchao Wang, Dacheng TaoICCV 2021 · 被引用 46 次
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma 等ICCV 2023 · 被引用 46 次
它引用的顶会 Paper9
- Factorizable Graph Convolutional NetworksYiding Yang, Zunlei Feng, Mingli Song, Xinchao WangNeurIPS 2020 · 被引用 175 次
- DGCN: Dynamic Graph Convolutional Network for Efficient Multi-Person Pose EstimationZhongwei Qiu, Kai Qiu, Jianlong Fu, Dongmei FuAAAI 2020 · 被引用 52 次
- 15 Keypoints Is All You NeedMichael Snower, Asim Kadav, Farley Lai, Hans Peter GrafCVPR 2020
- Distilling Knowledge From Graph Convolutional NetworksYiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao 等CVPR 2020
- Dynamic Graph Message Passing NetworksLi Zhang, Dan Xu, Anurag Arnab, Philip H. S. TorrCVPR 2020
相关 Paper
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu 等CVPR 2021
- Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosYu Cheng, Bo Wang, Bo Yang, Robby T. TanAAAI 2021 · 被引用 55 次
- GLoMOT: Efficient Online GNN-based Low-Frame-Rate Multi-Object TrackerYaxuan Hu, Jie Hua, Gang Wu, Yuhong Yang 等AAAI 2026
- CoMotion: Concurrent Multi-person 3D MotionAlejandro Newell, Peiyun Hu, Lahav Lipson, Stephan R. Richter 等ICLR 2025
