Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation
Zhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu, Yixing Gao, Yunjun Gao, Xiang Wang
Abstract
Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art methods strive to incorporate additional visual evidences from neighboring frames (supporting frames) to facilitate the pose estimation of the current frame (key frame). One aspect that has been obviated so far, is the fact that current methods directly aggregate unaligned contexts across frames. The spatial-misalignment between pose features of the current frame and neighboring frames might lead to unsatisfactory results. More importantly, existing approaches build upon the straightforward pose estimation loss, which unfortunately cannot constrain the network to fully leverage useful information from neighboring frames. To tackle these problems, we present a novel hierarchical alignment framework, which leverages coarse-to-fine deformations to progressively update a neighboring frame to align with the current frame at the feature level. We further propose to explicitly supervise the knowledge extraction from neighboring frames, guaranteeing that useful complementary cues are extracted. To achieve this goal, we theoretically analyzed the mutual information between the frames and arrived at a loss that maximizes the task-relevant mutual information. These allow us to rank No.1 in the Multi-frame Person Pose Estimation Challenge on benchmark dataset PoseTrack2017, and obtain state-of-the-art performance on benchmarks Sub-JHMDB and Pose-Track2018. Our code is released at https://github.com/Pose-Group/FAMI-Pose, hoping that it will be useful to the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d49c9995-b1a3-432f-b2eb-034112a56983Cited by top-tier papers26
- DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose EstimationRunyang Feng, Yixing Gao, Tze Ho Elden Tse, Xueqing Ma et al.ICCV 2023 · 46 citations
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture RecognitionYunan Li, Huizhou Chen, Guanwen Feng, Qiguang MiaoICCV 2023 · 10 citations
- Joint-Motion Mutual Learning for Pose Estimation in VideoSifan Wu, Haipeng Chen, Yifang Yin, Sihao Hu et al.ACM MM 2024 · 10 citations
Builds on21
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
- FaPN: Feature-aligned Pyramid Network for Dense Image PredictionShihua Huang, Zhichao Lu, Ran Cheng, Cheng HeICCV 2021 · 256 citations
- Indices Matter: Learning to Index for Deep Image MattingHao Lu, Yutong Dai, Chunhua Shen, Songcen XuICCV 2019 · 206 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
Related papers
- Deep Dual Consecutive Network for Human Pose EstimationZhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu et al.CVPR 2021
- Combining Detection and Tracking for Human Pose Estimation in VideosManchen Wang, Joseph Tighe, Davide ModoloCVPR 2020
- Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationSifan Wu, Haipeng Chen, Yingda Lyu, Shaojing Fan et al.AAAI 2026
- Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in VideoRunyang Feng, Yixing Gao, Xueqing Ma, Tze Ho Elden Tse et al.CVPR 2023
- Optimizing Human Pose Estimation Through Focused Human and Joint RegionsYingying Jiao, Zhigang Wang, Zhenguang Liu, Shaojing Fan et al.AAAI 2025 · 4 citations
