Reconstructing Hands in 3D with Transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, Jitendra Malik
Abstract
We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery, HaMeR, follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and robustness compared to previous work. The key to HaMeR's success lies in scaling up both the data used for training and the capacity of the deep network for hand reconstruction. For training data, we combine multiple datasets that contain 2D or 3D hand annotations. For the deep model, we use a large scale Vision Transformer architecture. Our final model consistently outperforms the previous baselines on popular 3D hand pose benchmarks. To further evaluate the effect of our design in non-controlled settings, we annotate existing in-the-wild datasets with 2D hand keypoint annotations. On this newly collected dataset of annotations, HInt, we demonstrate significant improvements over existing baselines. We will make our code, data and models publicly available upon publication. We make our code, data and models available on the project website: https://geopavlakos.github.io/hamer/. “It is because of his being armed with hands that man is the most intelligent animal.” Anaxagoras
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d59351d-e031-433b-a954-45c8ba5a9c86Cited by top-tier papers101
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoRyan Hoque, Peide Huang, David J. Yoon, Mouli Sivapurapu et al.ICLR 2026 · 248 citations
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng et al.ICML 2026 · 104 citations
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning MambaHaoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco et al.NeurIPS 2024 · 73 citations
Builds on8
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Towards A Richer 2D Understanding of Hands at ScaleTianyi Cheng, Dandan Shan, Ayda Hassen, Richard E. L. Higgins et al.NeurIPS 2023 · 48 citations
- DexYCB: A Benchmark for Capturing Hand Grasping of ObjectsYu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov et al.CVPR 2021
Related papers
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- End-to-End Human Pose and Mesh Reconstruction with TransformersKevin Lin, Lijuan Wang, Zicheng LiuCVPR 2021
- Reconstructing Humans with a Biomechanically Accurate SkeletonYan Xia, Xiaowei Zhou, Etienne Vouga, Qixing Huang et al.CVPR 2025
- Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesJunxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li et al.AAAI 2024 · 15 citations
- HORT: Monocular Hand-held Objects Reconstruction with TransformersZerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia SchmidICCV 2025 · 4 citations
