AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation
Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, Luan Tran, Cem Keskin
Abstract
We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging handobject interactions. The dataset includes synchronized egocentric and exocentric images sampled from the recent As-sembly101 dataset, in which participants assemble and disassemble take-apart toys. To obtain high-quality 3D hand pose annotations for the egocentric images, we develop an efficient pipeline, where we use an initial set of manual annotations to train a model to automatically annotate a much larger dataset. Our annotation model uses multi-view feature fusion and an iterative refinement scheme, and achieves an average keypoint error of 4.20 mm, which is 85% lower than the error of the original annotations in Assembly101. AssemblyHands provides 3.0M annotated images, including 490K egocentric images, making it the largest existing benchmark dataset for egocentric 3D hand pose estimation. Using this data, we develop a strong single-view baseline of 3D hand pose estimation from egocentric images. Furthermore, we design a novel action classification task to evaluate predicted 3D hand poses. Our study shows that having higher-quality hand poses directly improves the ability to recognize actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fed367cc-9b6d-4fce-80bc-8a841d7179ecCited by top-tier papers27
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa et al.CVPR 2024 · 110 citations
- ActSonic: Recognizing Everyday Activities from Inaudible Acoustic Wave Around the BodySaif Mahmud, Vineet Parikh, Qikang Liang, Ke Li et al.UbiComp 2025 · 24 citations
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song et al.ICML 2024 · 23 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the WildPatrick Rim, Kevin Harris, Braden Copple, Shangchen Han et al.CVPR 2026 · 5 citations
Builds on10
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
- MEgATrack: monochrome egocentric articulated hand-tracking for virtual realityShangchen Han, Beibei Liu, Randi Cabezas, Christopher D. Twigg et al.SIGGRAPH 2020 · 207 citations
Related papers
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- HOnnotate: A Method for 3D Annotation of Hand and Object PosesShreyas Hampali, Mahdi Rad, Markus Oberweger, Vincent LepetitCVPR 2020
- SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-trainingNie Lin, Takehiko Ohkawa, Yifei Huang, Mingfang Zhang et al.ICLR 2025
- Spectral Graphormer: Spectral Graph-based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color ImagesTze Ho Elden Tse, Franziska Mueller, Zhengyang Shen, Danhang Tang et al.ICCV 2023 · 13 citations
- Understanding Human Hands in Contact at Internet ScaleDandan Shan, Jiaqi Geng, Michelle Shu, David F. FouheyCVPR 2020
