THUNDR: Transformer-based 3D HUmaN Reconstruction with Markers
Mihai Zanfir, Andrei Zanfir, Eduard Gabriel Bazavan, William T. Freeman, Rahul Sukthankar, Cristian Sminchisescu
Abstract
We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and the regularizing, anthropometrically-preserving properties of a statistical human surface model like GHUM—a recently introduced, expressive full body statistical 3d human model, trained end-to-end. Our novel transformer-based prediction pipeline can focus on image regions relevant to the task, supports self-supervised regimes, and ensures that solutions are consistent with human anthropometry. We show state-of-the-art results on Human3.6M and 3DPW, for both the fully-supervised and the self-supervised models, for the task of inferring 3d human shape, joint positions, and global translation. Moreover, we observe very solid 3d reconstruction performance for difficult human poses collected in the wild.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9fff9f03-49f3-416e-843a-a40e89d9a1daCited by top-tier papers11
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- GOAL: Generating 4D Whole-Body Motion for Hand-Object GraspingOmid Taheri, Vasileios Choutas, Michael J. Black, Dimitrios TzionasCVPR 2022 · 103 citations
- Full-Body Motion from a Single Head-Mounted Device: Generating SMPL Poses from Partial ObservationsAndrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell et al.ICCV 2021 · 78 citations
- Neural Localizer Fields for Continuous 3D Human Pose and Shape EstimationIstván Sárándi, Gerard Pons-MollNeurIPS 2024 · 76 citations
- REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak SupervisionMihai Fieraru, Mihai Zanfir, Teodor Alexandru Szente, Eduard Gabriel Bazavan et al.NeurIPS 2021 · 44 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-CompareYuanlu Xu, Song-Chun Zhu, Tony TungICCV 2019 · 204 citations
- 3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image DataBenjamin Biggs, David Novotný, Sébastien Ehrhardt, Hanbyul Joo et al.NeurIPS 2020 · 79 citations
Related papers
- Neural Descent for Visual 3D Human Pose and ShapeAndrei Zanfir, Eduard Gabriel Bazavan, Mihai Zanfir, William T. Freeman et al.CVPR 2021
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- Deformable Mesh Transformer for 3D Human Mesh RecoveryYusuke YoshiyasuCVPR 2023
- HumanRAM: Feed-forward Human Reconstruction and Animation Model using TransformersZhiyuan Yu, Zhe Li, Hujun Bao, Can Yang et al.SIGGRAPH 2025 · 2 citations
