Cross-Domain 3D Hand Pose Estimation with Dual Modalities
Qiuxia Lin, Linlin Yang, Angela Yao
Abstract
Recent advances in hand pose estimation have shed light on utilizing synthetic data to train neural networks, which however inevitably hinders generalization to realworld data due to domain gaps. To solve this problem, we present a framework for cross-domain semi-supervised hand pose estimation and target the challenging scenario of learning models from labelled multi-modal synthetic data and unlabelled real-world data. To that end, we propose a dual-modality network that exploits synthetic RGB and synthetic depth images. For pre-training, our network uses multi-modal contrastive learning and attention-fused supervision to learn effective representations of the RGB images. We then integrate a novel self-distillation technique during fine-tuning to reduce pseudo-label noise. Experiments show that the proposed method significantly improves 3D hand pose estimation and 2D keypoint detection on benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e809b5f-c4e1-4b8d-bdc9-fe045fd77db0Cited by top-tier papers8
- Triangulation Residual Loss for Data-efficient 3D Pose EstimationJiachen Zhao, Tao Yu, Liang An, Yipeng Huang et al.NeurIPS 2023 · 13 citations
- Touchscreen-based Hand Tracking for Remote Whiteboard InteractionXinshuang Liu, Yizhong Zhang, Xin TongUIST 2024 · 8 citations
- SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationYinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing FuAAAI 2024 · 5 citations
- Synthetic-to-Real Pose Estimation with Geometric ReconstructionQiuxia Lin, Kerui Gu, Linlin Yang, Angela YaoNeurIPS 2023 · 4 citations
- Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose EstimationZhuoran Zhao, Linlin Yang, Pengzhan Sun, Pan Hui et al.CVPR 2025
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan et al.ICLR 2021 · 296 citations
Related papers
- SemiHand: Semi-supervised Hand Pose Estimation with ConsistencyLinlin Yang, Shicheng Chen, Angela YaoICCV 2021 · 42 citations
- Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior KnowledgeLong Zhao, Xi Peng, Yuxiao Chen, Mubbasir Kapadia et al.CVPR 2020
- Cross-Domain and Cross-Modal Knowledge Distillation in Domain Adaptation for 3D Semantic SegmentationMiaoyu Li, Yachao Zhang, Yuan Xie, Zuodong Gao et al.ACM MM 2022 · 30 citations
- Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ACM MM 2025 · 1 citation
- Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationXingyu Liu, Pengfei Ren, Yuanyuan Gao, Jingyu Wang et al.AAAI 2024 · 11 citations
