Joint-Motion Mutual Learning for Pose Estimation in Video
Sifan Wu, Haipeng Chen, Yifang Yin, Sihao Hu, Runyang Feng, Yingying Jiao, Ziqi Yang, Zhenguang Liu
摘要
Human pose estimation in videos has long been a compelling yet challenging task within the realm of computer vision. Nevertheless, this task remains difficult because of the complex video scenes, such as video defocus and self-occlusion. Recent methods strive to integrate multi-frame visual features generated by a backbone network for pose estimation. However, they often ignore the useful joint information encoded in the initial heatmap, which is a byproduct of the backbone generation. Comparatively, methods that attempt to refine the initial heatmap fail to consider any spatiotemporal motion features. As a result, the performance of existing methods for pose estimation falls short due to the lack of ability to leverage both local joint (heatmap) information and global motion (feature) dynamics.
To address this problem, we propose a novel joint-motion mutual learning framework for pose estimation, which effectively concentrates on both local joint dependency and global pixel-level motion dynamics. Specifically, we introduce a context-aware joint learner that adaptively leverages initial heatmaps and motion flow to retrieve robust local joint feature. Given that local joint feature
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WMamba: Wavelet-based Mamba for Face Forgery DetectionSiran Peng, Tianshuo Zhang, Li Gao, Xiangyu Zhu 等ACM MM 2025 · 被引用 17 次
- Two Heads Are Better than One: Distilling Large Language Model Features into Small Models with Feature Decomposition and MixtureTianhao Fu, Xinxin Xu, Weichen Xu, Jue Chen 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 被引用 1,030 次
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang 等ICCV 2021 · 被引用 363 次
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 被引用 360 次
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai 等ICCV 2023 · 被引用 150 次
相关 Paper
- SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled VideosYingying Jiao, Zhigang Wang, Sifan Wu, Shaojing Fan 等AAAI 2025 · 被引用 5 次
- Video-Based Human Pose Regression via Decoupled Space-Time AggregationJijie He, Wenwu YangCVPR 2024
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu 等CVPR 2021
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma 等ICCV 2023 · 被引用 46 次
- Multi-view 3D Smooth Human Pose Estimation based on Heatmap Filtering and Spatio-temporal InformationZehai Niu, Ke Lu, Jian Xue, Haifeng Ma 等ACM MM 2021 · 被引用 5 次
