Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent Tokens
Sen Yang, Wen Heng, Gang Liu, Guozhong Luo, Wankou Yang, Gang Yu
摘要
In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due to its inherent ambiguity. To improve precision, existing methods highly rely on the initialized mean pose and shape as prior estimates and parameter regression with an iterative error feedback manner. In addition, video-based approaches model the overall change over the image-level features to temporally enhance the single-frame feature, but fail to capture the rotational motion at the joint level, and cannot guarantee local temporal consistency. To address these issues, we propose a novel Transformer-based model with a design of independent tokens. First, we introduce three types of tokens independent of the image feature: joint rotation tokens, shape token, and camera token. By progressively interacting with image features through Transformer layers, these tokens learn to encode the prior knowledge of human 3D joint rotations, body shape, and position information from large-scale data, and are updated to estimate SMPL parameters conditioned on a given image. Second, benefiting from the proposed token-based representation, we further use a temporal model to focus on capturing the rotational temporal information of each joint, which is empirically conducive to preventing large jitters in local parts. Despite being conceptually simple, the proposed method attains superior performances on the 3DPW and Human3.6M datasets. Using ResNet-50 and Transformer architectures, it obtains 42.0 mm error on the PA-MPJPE metric of the challenging 3DPW, outperforming state-of-the-art counterparts by a large margin. Code will be publicly available at https://github.com/yangsenius/INT_HMR_Model
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PostureHMR: Posture Transformation for 3D Human Mesh RecoveryYu-Pei Song, Xiao Wu, Zhaoquan Yuanl, Jian-Jun Qiao 等CVPR 2024 · 被引用 11 次
- PoseGen: Learning to Generate 3D Human Pose Dataset with NeRFMohsen Gholami, Rabab Ward, Z. Jane WangAAAI 2024 · 被引用 3 次
- ARTS: Semi-Analytical Regressor using Disentangled Skeletal Representations for Human Mesh Recovery from VideosTao Tang, Hong Liu, Yingxuan You, Ti Wang 等ACM MM 2024 · 被引用 2 次
- SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive TokensChi Su, Xiaoxuan Ma, Jiajun Su, Yizhou WangCVPR 2025
- GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose EstimationJiyong Rao, Yu Wang, Shengjie ZhaoICLR 2026
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang 等CVPR 2022 · 被引用 403 次
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang 等ICCV 2021 · 被引用 648 次
- 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for TransformersYouze Xue, Jiansheng Chen, Yudong Zhang, Cheng Yu 等ACM MM 2022 · 被引用 10 次
- Global-to-Local Modeling for Video-Based 3D Human Pose and Shape EstimationXiaolong Shen, Zongxin Yang, Xiaohan Wang, Jianxin Ma 等CVPR 2023
