3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
Youze Xue, Jiansheng Chen, Yudong Zhang, Cheng Yu, Huimin Ma, Hongbing Ma
摘要
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- CDPNet: Cross-Modal Dual Phases Network for Point Cloud CompletionZhenjiang Du, Jiale Dou, Zhitao Liu, Jiwei Wei 等AAAI 2024 · 被引用 17 次
- Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryYongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang 等NeurIPS 2024 · 被引用 1 次
- WildAvatar: Learning In-the-wild 3D Avatars from the WebZihao Huang, Shoukang Hu, Guangcong Wang, Tianqi Liu 等CVPR 2025
相关 Paper
- Sampling is Matter: Point-Guided 3D Human Mesh ReconstructionJeonghwan Kim, Mi-Gyeong Gwon, Hyunwoo Park, Hyukmin Kwon 等CVPR 2023
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent TokensSen Yang, Wen Heng, Gang Liu, Guozhong Luo 等ICLR 2023 · 被引用 4 次
