TransPose: Keypoint Localization via Transformer
Sen Yang, Zhibin Quan, Mu Nie, Wankou Yang
Abstract
While CNN-based models have made remarkable progress on human pose estimation, what spatial dependencies they capture to localize keypoints remains unclear. In this work, we propose a model called Trans-Pose, which introduces Transformer for human pose estimation. The attention layers built in Transformer enable our model to capture long-range relationships efficiently and also can reveal what dependencies the predicted key-points rely on. To predict keypoint heatmaps, the last attention layer acts as an aggregator, which collects contributions from image clues and forms maximum positions of keypoints. Such a heatmap-based localization approach via Transformer conforms to the principle of Activation Maximization [19]. And the revealed dependencies are image-specific and fine-grained, which also can provide evidence of how the model handles special cases, e.g., occlusion. The experiments show that TransPose achieves 75.8 AP and 75.0 AP on COCO validation and test-dev sets, while being more lightweight and faster than mainstream CNN architectures. The TransPose model also transfers very well on MPII benchmark, achieving superior performance on the test set when fine-tuned with small training costs. Code and pre-trained models are publicly available1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9841102f-77bd-4934-a24f-3ca3858e3fffCited by top-tier papers42
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- TransWeather: Transformer-based Restoration of Images Degraded by Adverse Weather ConditionsJeya Maria Jose Valanarasu, Rajeev Yasarla, Vishal M. PatelCVPR 2022 · 350 citations
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionYuanhao Cai, Jing Lin, Xiaowan Hu, Haoqian Wang et al.CVPR 2022 · 310 citations
- Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingYuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan et al.NeurIPS 2022 · 222 citations
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi et al.CVPR 2022 · 116 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- How much Position Information Do Convolutional Neural Networks Encode?Md. Amirul Islam, Sen Jia, Neil D. B. BruceICLR 2020 · 392 citations
Related papers
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 155 citations
- Location-Free Human Pose EstimationXixia Xu, Yingguo Gao, Ke Yan, Xue Lin et al.CVPR 2022 · 13 citations
- SHaRPose: Sparse High-Resolution Representation for Human Pose EstimationXiaoqi An, Lin Zhao, Chen Gong, Nannan Wang et al.AAAI 2024 · 36 citations
- End-to-End Multi-Person Pose Estimation with TransformersDahu Shi, Xing Wei, Liangqi Li, Ye Ren et al.CVPR 2022 · 147 citations
