Combining Detection and Tracking for Human Pose Estimation in Videos
Manchen Wang, Joseph Tighe, Davide Modolo
Abstract
We propose a novel top-down approach that tackles the problem of multi-person human pose estimation and tracking in videos. In contrast to existing top-down approaches, our method is not limited by the performance of its person detector and can predict the poses of person instances not localized. It achieves this capability by propagating known person locations forward and backward in time and searching for poses in those regions. Our approach consists of three components: (i) a Clip Tracking Network that performs body joint detection and tracking simultaneously on small video clips; (ii) a Video Tracking Pipeline that merges the fixed-length tracklets produced by the Clip Tracking Network to arbitrary length tracks; and (iii) a Spatial-Temporal Merging procedure that refines the joint locations based on spatial and temporal smoothing terms. Thanks to the precision of our Clip Tracking Network and our merging procedure, our approach produces very accurate joint predictions and can fix common mistakes on hard scenarios like heavily entangled people. Our approach achieves state-of-the-art results on both joint detection and tracking, on both the PoseTrack 2017 and 2018 datasets, and against all top-down and bottom-down approaches. We propose a novel top-down approach that overcomes these problems and enables us to reap the benefits of top down methods for multi-person pose estimation in videos. We detect person bounding boxes on each frame and then
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe965a1e-bd56-4320-9925-e1d71956fff0Cited by top-tier papers25
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu et al.CVPR 2022 · 76 citations
- Test-Time Personalization with a Transformer for Human Pose EstimationYizhuo Li, Miao Hao, Zonglin Di, Nitesh B. Gundavarapu et al.NeurIPS 2021 · 58 citations
- PoseTrack21: A Dataset for Person Search, Multi-Object Tracking and Multi-Person Pose TrackingAndreas Doering, Di Chen, Shanshan Zhang, Bernt Schiele et al.CVPR 2022 · 47 citations
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
- Efficient Video Instance Segmentation via Tracklet Query and ProposalJialian Wu, Sudhir Yarram, Hui Liang, Tian Lan et al.CVPR 2022 · 33 citations
Builds on1
Related papers
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 2 citations
- Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up NetworksYu Cheng, Bo Wang, Bo Yang, Robby T. TanCVPR 2021
- Learning Dynamics via Graph Neural Networks for Human Pose Estimation and TrackingYiding Yang, Zhou Ren, Haoxiang Li, Chunluan Zhou et al.CVPR 2021
- Video-Based Human Pose Regression via Decoupled Space-Time AggregationJijie He, Wenwu YangCVPR 2024
