Joint-Motion Mutual Learning for Pose Estimation in Video
Sifan Wu, Haipeng Chen, Yifang Yin, Sihao Hu, Runyang Feng, Yingying Jiao, Ziqi Yang, Zhenguang Liu
Abstract
Human pose estimation in videos has long been a compelling yet challenging task within the realm of computer vision. Nevertheless, this task remains difficult because of the complex video scenes, such as video defocus and self-occlusion. Recent methods strive to integrate multi-frame visual features generated by a backbone network for pose estimation. However, they often ignore the useful joint information encoded in the initial heatmap, which is a byproduct of the backbone generation. Comparatively, methods that attempt to refine the initial heatmap fail to consider any spatiotemporal motion features. As a result, the performance of existing methods for pose estimation falls short due to the lack of ability to leverage both local joint (heatmap) information and global motion (feature) dynamics.
To address this problem, we propose a novel joint-motion mutual learning framework for pose estimation, which effectively concentrates on both local joint dependency and global pixel-level motion dynamics. Specifically, we introduce a context-aware joint learner that adaptively leverages initial heatmaps and motion flow to retrieve robust local joint feature. Given that local joint feature
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc40b24b-ba33-4860-bb0d-c11ce2e848d1Cited by top-tier papers2
- WMamba: Wavelet-based Mamba for Face Forgery DetectionSiran Peng, Tianshuo Zhang, Li Gao, Xiangyu Zhu et al.ACM MM 2025 · 17 citations
- Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and MixtureTianhao Fu, Xinxin Xu, Weichen Xu, Jue Chen et al.AAAI 2026 · 2 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 360 citations
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai et al.ICCV 2023 · 150 citations
Related papers
- SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosYingying Jiao, Zhigang Wang, Sifan Wu, Shaojing Fan et al.AAAI 2025 · 5 citations
- Video-Based Human Pose Regression via Decoupled Space-Time AggregationJijie He, Wenwu YangCVPR 2024
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
- Multi-view 3D Smooth Human Pose Estimation based on Heatmap Filtering and Spatio-temporal InformationZehai Niu, Ke Lu, Jian Xue, Haifeng Ma et al.ACM MM 2021 · 5 citations
