3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
Youze Xue, Jiansheng Chen, Yudong Zhang, Cheng Yu, Huimin Ma, Hongbing Ma
Abstract
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 88178535-abe8-4bfc-afc5-a7ff63ab893cCited by top-tier papers3
- CDPNet: Cross-Modal Dual Phases Network for Point Cloud CompletionZhenjiang Du, Jiale Dou, Zhitao Liu, Jiwei Wei et al.AAAI 2024 · 17 citations
- Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryYongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang et al.NeurIPS 2024 · 1 citation
- WildAvatar: Learning In-the-wild 3D Avatars from the WebZihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu et al.CVPR 2025
Related papers
- Sampling is Matter: Point-Guided 3D Human Mesh ReconstructionJeonghwan Kim, Mi-Gyeong Gwon, Hyunwoo Park, Hyukmin Kwon et al.CVPR 2023
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent TokensSen Yang, Wen Heng, Gang Liu, Guozhong Luo et al.ICLR 2023 · 4 citations
