TokenPose: Learning Keypoint Tokens for Human Pose Estimation
Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, Erjin Zhou
摘要
Human pose estimation deeply relies on visual clues and anatomical constraints between parts to locate keypoints. Most existing CNN-based methods do well in visual representation, however, lacking in the ability to explicitly learn the constraint relationships between keypoints. In this paper, we propose a novel approach based on Token representation for human Pose estimation (TokenPose). In detail, each keypoint is explicitly embedded as a token to simultaneously learn constraint relationships and appearance cues from images. Extensive experiments show that the small and large TokenPose models are on par with state-of-the-art CNN-based counterparts while being more lightweight. Specifically, our TokenPose-S and TokenPose-L achieve 72.5 AP and 75.8 AP on COCO validation dataset respectively, with significant reduction in parameters (↓ 80.6% ; ↓ 56.8%) and GFLOPs (↓ 75.3%; ↓ 24.7%). Code is publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingYuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan 等NeurIPS 2022 · 被引用 222 次
- Generating Transferable 3D Adversarial Point Cloud via Random Perturbation FactorizationBangyan He, Jian Liu, Yiming Li, Siyuan Liang 等AAAI 2023 · 被引用 53 次
- SHaRPose: Sparse High-Resolution Representation for Human Pose EstimationXiaoqi An, Lin Zhao, Chen Gong, Nannan Wang 等AAAI 2024 · 被引用 36 次
- Rethinking pose estimation in crowds: overcoming the detection information bottleneck and ambiguityMu Zhou, Lucas Stoffl, Mackenzie Weygandt Mathis, Alexander MathisICCV 2023 · 被引用 28 次
- Lightweight Super-Resolution Head for Human Pose EstimationHaonan Wang, Jie Liu, Jie Tang, Gangshan WuACM MM 2023 · 被引用 23 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- Human Pose as Compositional TokensZigang Geng, Chunyu Wang, Yixuan Wei, Ze Liu 等CVPR 2023
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 被引用 360 次
- DistilPose: Tokenized Pose Regression with Heatmap DistillationSuhang Ye, Yingyi Zhang, Jie Hu, Liujuan Cao 等CVPR 2023
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla 等NeurIPS 2024 · 被引用 4 次
- Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose EstimationJia Li, Wen Su, Zengfu WangAAAI 2020 · 被引用 104 次
