Pose Recognition With Cascade Transformers
Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, Zhuowen Tu
Abstract
In this paper, we present a regression-based pose recognition method using cascade Transformers. One way to categorize the existing approaches in this domain is to separate them into 1). heatmap-based and 2). regressionbased. In general, heatmap-based methods achieve higher accuracy but are subject to various heuristic designs (not end-to-end mostly), whereas regression-based approaches attain relatively lower accuracy but they have less intermediate non-differentiable steps. Here we utilize the encoderdecoder structure in Transformers to perform regressionbased person and keypoint detection that is general-purpose and requires less heuristic design compared with the existing approaches. We demonstrate the keypoint hypothesis (query) refinement process across different self-attention layers to reveal the recursive self-attention mechanism in Transformers. In the experiments, we report competitive results for pose recognition when compared with the competing regression-based methods. * indicates equal contribution. Code: https://github.com/mlpc-ucsd/PRTR . Work performed during internships of K. Li, S.Wang, and X. Zhang with UC San Diego.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d42f9c5-5248-4eab-a70f-2ce27b267f9dCited by top-tier papers44
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionYuanhao Cai, Jing Lin, Xiaowan Hu, Haoqian Wang et al.CVPR 2022 · 310 citations
- EDTER: Edge Detection with TransformerMengyang Pu, Yaping Huang, Yuming Liu, Qingji Guan et al.CVPR 2022 · 224 citations
- Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive ImagingYuanhao Cai, Jing Lin, Haoqian Wang, Xin Yuan et al.NeurIPS 2022 · 222 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Single-Stage Multi-Person Pose MachinesXuecheng Nie, Jiashi Feng, Jianfeng Zhang, Shuicheng YanICCV 2019 · 246 citations
- Distribution-Aware Coordinate Representation for Human Pose EstimationFeng Zhang, Xiatian Zhu, Hanbin Dai, Mao Ye et al.CVPR 2020
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi et al.CVPR 2020
- The Devil Is in the Details: Delving Into Unbiased Data Processing for Human Pose EstimationJunjie Huang, Zheng Zhu, Feng Guo, Guan HuangCVPR 2020
Related papers
- Towards Accurate Facial Landmark Detection via Cascaded TransformersHui Li, Zidong Guo, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 45 citations
- Group Pose: A Simple Baseline for End-to-End Multi-person Pose EstimationHuan Liu, Qiang Chen, Zichang Tan, Jiang-Jiang Liu et al.ICCV 2023 · 50 citations
- TransPose: Keypoint Localization via TransformerSen Yang, Zhibin Quan, Mu Nie, Wankou YangICCV 2021 · 360 citations
- HAT: Hierarchical Aggregation Transformers for Person Re-identificationGuowen Zhang, Pingping Zhang, Jinqing Qi, Huchuan LuACM MM 2021 · 159 citations
- Cascade Transformers for End-to-End Person SearchRui Yu, Dawei Du, Rodney LaLonde, Daniel Davila et al.CVPR 2022 · 86 citations
