Aligning Latent Spaces for 3D Hand Pose Estimation
Linlin Yang, Shile Li, Dongheui Lee, Angela Yao
Abstract
Hand pose estimation from monocular RGB inputs is a highly challenging task. Many previous works for monocular settings only used RGB information for training despite the availability of corresponding data in other modalities such as depth maps. In this work, we propose to learn a joint latent representation that leverages other modalities as weak labels to boost the RGB-based hand pose estimator. By design, our architecture is highly flexible in embedding various diverse modalities such as heat maps, depth maps and point clouds. In particular, we find that encoding and decoding the point cloud of the hand surface can improve the quality of the joint latent representation. Experiments show that with the aid of other modalities during training, our proposed method boosts the accuracy of RGB-based hand pose estimation systems and significantly outperforms state-of-the-art on two public benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15ce27ea-5fee-459a-b6d3-b9b034a7d695Cited by top-tier papers28
- MobRecon: Mobile-Friendly Hand Mesh Reconstruction from Monocular ImageXingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang et al.CVPR 2022 · 97 citations
- Towards Accurate Alignment in Real-time 3D Hand-Mesh ReconstructionXiao Tang, Tianyu Wang, Chi-Wing FuICCV 2021 · 83 citations
- EventHands: Real-Time Neural 3D Hand Pose Estimation from an Event StreamViktor Rudnev, Vladislav Golyanik, Jiayi Wang, Hans-Peter Seidel et al.ICCV 2021 · 66 citations
- Hand Image Understanding via Deep Multi-Task LearningXiong Zhang, Hongsheng Huang, Jianchao Tan, Hongmin Xu et al.ICCV 2021 · 66 citations
- Lightweight Multi-person Total Motion Capture Using Sparse Multi-view CamerasYuxiang Zhang, Zhe Li, Liang An, Mengcheng Li et al.ICCV 2021 · 47 citations
Related papers
- Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationXingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang et al.AAAI 2024 · 11 citations
- HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth MapJameel Malik, Ibrahim Abdelaziz, Ahmed Elhayek, Soshi Shimada et al.CVPR 2020
- Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior KnowledgeLong Zhao, Xi Peng, Yuxiao Chen, Mubbasir Kapadia et al.CVPR 2020
- MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose SynthesisZhenyu Wu, Duc Hoang, Shih-Yao Lin, Yusheng Xie et al.ACM MM 2020 · 16 citations
- Cross-Domain 3D Hand Pose Estimation with Dual ModalitiesQiuxia Lin, Linlin Yang, Angela YaoCVPR 2023
