Deep Dual Consecutive Network for Human Pose Estimation
Zhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu, Shouling Ji, Bailin Yang, Xun Wang
Abstract
Multi-frame human pose estimation in complicated situations is challenging. Although state-of-the-art human joints detectors have demonstrated remarkable results for static images, their performances come short when we apply these models to video sequences. Prevalent shortcomings include the failure to handle motion blur, video defocus, or pose occlusions, arising from the inability in capturing the temporal dependency among video frames. On the other hand, directly employing conventional recurrent neural networks incurs empirical difficulties in modeling spatial contexts, especially for dealing with pose occlusions. In this paper, we propose a novel multi-frame human pose estimation framework, leveraging abundant temporal cues between video frames to facilitate keypoint detection. Three modular components are designed in our framework. A Pose Temporal Merger encodes keypoint spatiotemporal context to generate effective searching scopes while a Pose Residual Fusion module computes weighted pose residuals in dual directions. These are then processed via our Pose Correction Network for efficient refining of pose estimations. Our method ranks No.1 in the Multi-frame Person Pose Estimation Challenge on the large-scale benchmark datasets PoseTrack2017 and PoseTrack2018. We have released our code, hoping to inspire future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68e0c2e7-4c7d-49e0-8e68-0d871ce7b6a3Cited by top-tier papers32
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu et al.CVPR 2022 · 76 citations
- Motion Prediction using Trajectory CuesZhenguang Liu, Pengxiang Su, Shuang Wu, Xuanjing Shen et al.ICCV 2021 · 63 citations
- Locate and Verify: A Two-Stream Network for Improved Deepfake DetectionChao Shuai, Jieming Zhong, Shuang Wu, Feng Lin et al.ACM MM 2023 · 52 citations
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
- Self-Regulation for Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.ICCV 2021 · 42 citations
Builds on8
- Aggregated Multi-GANs for Controlled 3D Human Motion PredictionZhenguang Liu, Kedi Lyu, Shuang Wu, Haipeng Chen et al.AAAI 2021 · 64 citations
- Mixture Dense Regression for Object Detection and Human Pose EstimationAli Varamesh, Tinne TuytelaarsCVPR 2020
- Distribution-Aware Coordinate Representation for Human Pose EstimationFeng Zhang, Xiatian Zhu, Hanbin Dai, Mao Ye et al.CVPR 2020
- 15 Keypoints Is All You NeedMichael Snower, Asim Kadav, Farley Lai, Hans Peter GrafCVPR 2020
- UniPose: Unified Human Pose Estimation in Single Images and VideosBruno Artacho, Andreas E. SavakisCVPR 2020
Related papers
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Video-Based Human Pose Regression via Decoupled Space-Time AggregationJijie He, Wenwu YangCVPR 2024
- Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
- End-to-End Multi-Person Pose Estimation with Pose-Aware Video TransformerYonghui Yu, Jiahang Cai, Xun Wang, Wenwu YangAAAI 2026 · 2 citations
- Learning Dynamics via Graph Neural Networks for Human Pose Estimation and TrackingYiding Yang, Zhou Ren, Haoxiang Li, Chunluan Zhou et al.CVPR 2021
