The Devil Is in the Details: Delving Into Unbiased Data Processing for Human Pose Estimation
Junjie Huang, Zheng Zhu, Feng Guo, Guan Huang
摘要
Being a fundamental component in training and inference, data processing has not been systematically considered in human pose estimation community, to the best of our knowledge. In this paper, we focus on this problem and find that the devil of human pose estimation evolution is in the biased data processing. Specifically, by investigating the standard data processing in state-of-the-art approaches mainly including coordinate system transformation and keypoint format transformation (i.e., encoding and decoding), we find that the results obtained by common flipping strategy are unaligned with the original ones in inference. Moreover, there is a statistical error in some keypoint format transformation methods. Two problems couple together, significantly degrade the pose estimation performance and thus lay a trap for the research community. This trap has given bone to many suboptimal remedies, which are always unreported, confusing but influential. By causing failure in reproduction and unfair in comparison, the unreported remedies seriously impedes the technological development. To tackle this dilemma from the source, we propose Unbiased Data Processing (UDP) consist of two technique aspect for the two aforementioned problems respectively (i.e., unbiased coordinate system transformation and unbiased keypoint format transformation). Base on UDP, we wipe out the trap by giving out a deep insight of the existing biased data processing pipeline, whose origin, effects and some confusing remedies are thoroughly studied. Besides, as a model-agnostic approach and a superior solution, UDP successfully pushes the performance boundary of human pose estimation. For example on COCO test-dev set, UDP promotes top-down method HRNet-W32-256×192 by 1.7 AP (73.5 to 75.2) for free and promotes bottom-up methods HRNet-W32-512×512 by 2.7 AP with an acceleration of 6.1 times. The HRNet-W48-384×288 equipped with UDP achieves 76.5 AP and sets a new state-of-the-art for human pose estimation. As a meaningful milestone for pursuing high performance human pose estimation, UDP has been the key base of the winner in 2020 COCO Keypoint Detection Challenge. The code is public available for reference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin 等NeurIPS 2021 · 被引用 357 次
- Multi-Instance Pose Networks: Rethinking Top-Down Pose EstimationRawal Khirodkar, Visesh Chari, Amit Agrawal, Ambrish TyagiICCV 2021 · 被引用 80 次
- Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose EstimationZhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 等CVPR 2022 · 被引用 76 次
- The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose EstimationGuillem Brasó, Nikita Kister, Laura Leal-TaixéICCV 2021 · 被引用 50 次
它引用的顶会 Paper3
- Single-Stage Multi-Person Pose MachinesXuecheng Nie, Jiashi Feng, Jianfeng Zhang, Shuicheng YanICCV 2019 · 被引用 246 次
- Distribution-Aware Coordinate Representation for Human Pose EstimationFeng Zhang, Xiatian Zhu, Hanbin Dai, Mao Ye 等CVPR 2020
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi 等CVPR 2020
相关 Paper
- Implicit Decouple Network for Efficient Pose EstimationLei Zhao, Le Han, Min Yao, Nenggan ZhengACM MM 2023 · 被引用 1 次
- Rethinking the Heatmap Regression for Bottom-Up Human Pose EstimationZhengxiong Luo, Zhicheng Wang, Yan Huang, Liang Wang 等CVPR 2021
- Bottom-Up Human Pose Estimation via Disentangled Keypoint RegressionZigang Geng, Ke Sun, Bin Xiao, Zhaoxiang Zhang 等CVPR 2021
- Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose EstimationXiaoke Jiang, Donghai Li, Hao Chen, Ye Zheng 等CVPR 2022 · 被引用 54 次
- Semantic-aware Transfer with Instance-adaptive Parsing for Crowded Scenes Pose EstimationXuanhan Wang, Lianli Gao, Yan Dai, Yixuan Zhou 等ACM MM 2021 · 被引用 14 次
