Lightweight Super-Resolution Head for Human Pose Estimation
Haonan Wang, Jie Liu, Jie Tang, Gangshan Wu
Abstract
Heatmap-based methods have become the mainstream method for pose estimation due to their superior performance. However, heatmapbased approaches suffer from significant quantization errors with downscale heatmaps, which result in limited performance and the detrimental effects of intermediate supervision. Previous heatmapbased methods relied heavily on additional post-processing to mitigate quantization errors. Some heatmap-based approaches improve the resolution of feature maps by using multiple costly upsampling layers to improve localization precision. To solve the above issues, we creatively view the backbone network as a degradation process and thus reformulate the heatmap prediction as a Super-Resolution (SR) task. We first propose the SR head, which predicts heatmaps with a spatial resolution higher than the input feature maps (or even consistent with the input image) by super-resolution, to effectively reduce the quantization error and the dependence on further post-processing. Besides, we propose SRPose to gradually recover the HR heatmaps from LR heatmaps and degraded features in a coarse-to-fine manner. To reduce the training difficulty of HR heatmaps, SRPose applies SR heads to supervise the intermediate features in each stage. In addition, the SR head is a lightweight and generic head that applies to top-down and bottom-up methods. Extensive experiments on the COCO, MPII, and Crowd-Pose datasets show that SRPose outperforms the corresponding heatmap-based approaches. The code and models are available at https://github.com/haonanwang0522/SRPose.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09e7ee6b-3ae8-4231-8853-0412d95aa47dBuilds on15
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 360 citations
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang et al.ICCV 2021 · 286 citations
Related papers
- Continuous Heatmap Regression for Pose Estimation via Implicit Neural RepresentationShengxiang Hu, Huaijiang Sun, Dong Wei, Xiaoning Sun et al.NeurIPS 2024 · 5 citations
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi et al.CVPR 2020
- A Characteristic Function-Based Method for Bottom-Up Human Pose EstimationHaoxuan Qu, Yujun Cai, Lin Geng Foo, Ajay Kumar et al.CVPR 2023
- Rethinking the Heatmap Regression for Bottom-Up Human Pose EstimationZhengxiong Luo, Zhicheng Wang, Yan Huang, Liang Wang et al.CVPR 2021
- Distribution-Aware Coordinate Representation for Human Pose EstimationFeng Zhang, Xiatian Zhu, Hanbin Dai, Mao Ye et al.CVPR 2020
