SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation
Xiaoqi An, Lin Zhao, Chen Gong, Nannan Wang, Di Wang, Jian Yang
摘要
High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational burden. In this paper, we address the following question: "Only sparse human keypoint locations are detected for human pose estimation, is it really necessary to describe the whole image in a dense, high-resolution manner?" Based on dynamic transformer models, we propose a framework that only uses Sparse High-resolution Representations for human Pose estimation (SHaRPose). In detail, SHaRPose consists of two stages. At the coarse stage, the relations between image regions and keypoints are dynamically mined while a coarse estimation is generated. Then, a quality predictor is applied to decide whether the coarse estimation results should be refined. At the fine stage, SHaRPose builds sparse high-resolution representations only on the regions related to the keypoints and provides refined high-precision human pose estimations. Extensive experiments demonstrate the outstanding performance of the proposed method. Specifically, compared to the state-of-the-art method ViTPose, our model SHaRPose-Base achieves 77.4 AP (+0.5 AP) on the COCO validation set and 76.7 AP (+0.5 AP) on the COCO test-dev set, and infers at a speed of 1.4x faster than ViTPose-Base. Code is available at https://github.com/AnxQ/sharpose.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Exploiting Multimodal Spatial-temporal Patterns for Video Object TrackingXiantao Hu, Ying Tai, Xu Zhao, Chen Zhao 等AAAI 2025 · 被引用 65 次
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等AAAI 2025 · 被引用 8 次
- Faster Vision Transformers with Adaptive PatchesRohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang 等ICLR 2026 · 被引用 8 次
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion TransformersDahye Kim, Deepti Ghadiyaram, Raghudeep GaddeCVPR 2026 · 被引用 3 次
- Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose EstimationXiaoqi An, Lin Zhao, Chen Gong, Jun Li 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li 等CVPR 2022 · 被引用 835 次
相关 Paper
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 被引用 360 次
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang 等ICCV 2021 · 被引用 363 次
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi 等CVPR 2020
- Bottom-Up Human Pose Estimation via Disentangled Keypoint RegressionZigang Geng, Ke Sun, Bin Xiao, Zhaoxiang Zhang 等CVPR 2021
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
