WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, Stefanos Zafeiriou
Abstract
Figure 1. We propose WiLoR, a full-stack in-the-Wild Localization and 3D hand Reconstruction method. WiLoR first localizes and defines the handedness of the detected hands which are then lifted to 3D using a transformer-based hand pose estimation module. To aid high-fidelity reconstructions and facilitate image-alignment, we introduce a refinement module that extracts localized features to correct misaligned poses. WiLoR achieves state-of-the-art performance under different benchmark datasets while boosting the temporal coherence of image-based 3D hand pose estimation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfc3ddd7-40fe-44cd-8cb7-99a9b4d388baCited by top-tier papers36
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human VideosGu Zhang, Qicheng Xu, Haozhe Zhang, Jianhan Ma et al.CVPR 2026 · 23 citations
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti et al.CVPR 2026 · 16 citations
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan et al.CVPR 2026 · 15 citations
- Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language GeneratorRonglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng et al.ICCV 2025 · 9 citations
Builds on41
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- HORT: Monocular Hand-held Objects Reconstruction with TransformersZerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia SchmidICCV 2025 · 4 citations
- HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose EstimationLin Huang, Jianchao Tan, Jingjing Meng, Ji Liu et al.ACM MM 2020 · 59 citations
- MeMaHand: Exploiting Mesh-Mano Interaction for Single Image Two-Hand ReconstructionCongyi Wang, Feida Zhu, Shilei WenCVPR 2023
- A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB ImageChanglong Jiang, Yang Xiao, Cunlin Wu, Mingyang Zhang et al.CVPR 2023
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric VideosJinglei Zhang, Jiankang Deng, Chao Ma, Rolandos Alexandros PotamiasCVPR 2025
