Human Pose as Compositional Tokens
Zigang Geng, Chunyu Wang, Yixuan Wei, Ze Liu, Houqiang Li, Han Hu
Abstract
Human pose is typically represented by a coordinate vector of body joints or their heatmap embeddings. While easy for data processing, unrealistic pose estimates are admitted due to the lack of dependency modeling between the body joints. In this paper, we present a structured representation, named Pose as Compositional Tokens (PCT), to explore the joint dependency. It represents a pose by M discrete tokens with each characterizing a sub-structure with several interdependent joints (see Figure 1 ). The compositional design enables it to achieve a small reconstruction error at a low cost. Then we cast pose estimation as a classification task. In particular, we learn a classifier to predict the categories of the M tokens from an image. A prelearned decoder network is used to recover the pose from the tokens without further post-processing. We show that it achieves better or comparable pose estimation results as the existing methods in general scenarios, yet continues to work well when occlusion occurs, which is ubiquitous in practice. The code and models are publicly available at https://github.com/Gengzigang/PCT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22ec4ee6-89da-44d5-9ef6-9d3eb1c31015Cited by top-tier papers23
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong et al.NeurIPS 2024 · 51 citations
- MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete RepresentationsHeyuan Yao, Zhenhua Song, Yuyang Zhou, Tenglong Ao et al.SIGGRAPH 2024 · 34 citations
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
- WalkTheDog: Cross-Morphology Motion Alignment via Phase ManifoldsPeizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-HornungSIGGRAPH 2024 · 10 citations
- : Discrete Diffusion Model for Occluded 3D Human Pose EstimationWeiquan Wang, Jun Xiao, Chunping Wang, Wei Liu et al.NeurIPS 2024 · 4 citations
Builds on36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
Related papers
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang et al.ICCV 2021 · 363 citations
- Object Recognition as Next Token PredictionKaiyu Yue, Bor-Chun Chen, Jonas Geiping, Hengduo Li et al.CVPR 2024 · 6 citations
- Causal-Inspired Multitask Learning for Video-Based Human Pose EstimationHaipeng Chen, Sifan Wu, Zhigang Wang, Yifang Yin et al.AAAI 2025 · 7 citations
- Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationWenhao Li, Mengyuan Liu, Hong Liu, Pichao Wang et al.CVPR 2024
- Progressive Bi-C3D Pose Grammar for Human Pose EstimationLu Zhou, Yingying Chen, Jinqiao Wang, Hanqing LuAAAI 2020 · 5 citations
