Human Pose as Compositional Tokens
Zigang Geng, Chunyu Wang, Yixuan Wei, Ze Liu, Houqiang Li, Han Hu
摘要
Human pose is typically represented by a coordinate vector of body joints or their heatmap embeddings. While easy for data processing, unrealistic pose estimates are admitted due to the lack of dependency modeling between the body joints. In this paper, we present a structured representation, named Pose as Compositional Tokens (PCT), to explore the joint dependency. It represents a pose by M discrete tokens with each characterizing a sub-structure with several interdependent joints (see Figure 1 ). The compositional design enables it to achieve a small reconstruction error at a low cost. Then we cast pose estimation as a classification task. In particular, we learn a classifier to predict the categories of the M tokens from an image. A prelearned decoder network is used to recover the pose from the tokens without further post-processing. We show that it achieves better or comparable pose estimation results as the existing methods in general scenarios, yet continues to work well when occlusion occurs, which is ubiquitous in practice. The code and models are publicly available at https://github.com/Gengzigang/PCT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- MoGenTS: Motion Generation based on Spatial-Temporal Joint ModelingWeihao Yuan, Yisheng He, Weichao Shen, Yuan Dong 等NeurIPS 2024 · 被引用 51 次
- MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete RepresentationsHeyuan Yao, Zhenhua Song, Yuyang Zhou, Tenglong Ao 等SIGGRAPH 2024 · 被引用 34 次
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou 等AAAI 2024 · 被引用 14 次
- WalkTheDog: Cross-Morphology Motion Alignment via Phase ManifoldsPeizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-HornungSIGGRAPH 2024 · 被引用 10 次
- : Discrete Diffusion Model for Occluded 3D Human Pose EstimationWeiquan Wang, Jun Xiao, Chunping Wang, Wei Liu 等NeurIPS 2024 · 被引用 4 次
它引用的顶会 Paper36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
相关 Paper
- TokenPose: Learning Keypoint Tokens for Human Pose EstimationYanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang 等ICCV 2021 · 被引用 363 次
- Object Recognition as Next Token PredictionKaiyu Yue, Bor-Chun Chen, Jonas Geiping, Hengduo Li 等CVPR 2024 · 被引用 6 次
- Causal-Inspired Multitask Learning for Video-Based Human Pose EstimationHaipeng Chen, Sifan Wu, Zhigang Wang, Yifang Yin 等AAAI 2025 · 被引用 7 次
- Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose EstimationWenhao Li, Mengyuan Liu, Hong Liu, Pichao Wang 等CVPR 2024
- Progressive Bi-C3D Pose Grammar for Human Pose EstimationLu Zhou, Yingying Chen, Jinqiao Wang, Hanqing LuAAAI 2020 · 被引用 5 次
