UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
Tianhao Han, HaoYang ZHANG, Liang Xie, Haochen Chang, Kun Gao, Yuan Cheng, Pengfei Ren, Erwei Yin
Abstract
Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or multiview consistency constraints, as the driving force to optimize networks and progressively refine pose accuracy. However, these methods are highly susceptible to noisy pseudo-labels and overlook the importance of fully exploiting fine-grained spatial correlations, which undermines the stability of model training. To address these issues, we propose UST-Hand, a self-supervised learning framework that estimates uncertainty distribution of hand pose and constructs a probabilistic point cloud feature space, which enables the complex spatiotemporal relationship modeling. UST-Hand employs a conditional normalizing flow model to capture hand pose distributions and samples diversity hypotheses, facilitating robust learning under noisy pseudo-labels supervision with enhanced stability. These multi-hypothesis are mapped to a unified probabilistic 3D point cloud space for multiview and temporal feature interaction, comprehensively exploring hand motion patterns and fine-grained spatial correlations. Extensive experiments on three challenging datasets demonstrate that UST-Hand achieves state-of-the-art performance, outperforming existing self-supervised methods by up to 37.8% in Mean Per Vertex Position Error (MPVPE).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5c6de52-6b1e-4686-8eff-a80d6da842feBuilds on38
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose EstimationWenhao Li, Hong Liu, Hao Tang, Pichao Wang et al.CVPR 2022 · 403 citations
- Human Pose Regression with Residual Log-likelihood EstimationJiefeng Li, Siyuan Bian, Ailing Zeng, Can Wang et al.ICCV 2021 · 286 citations
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
Related papers
- Model-Based 3D Hand Reconstruction via Self-Supervised LearningYujin Chen, Zhigang Tu, Di Kang, Linchao Bao et al.CVPR 2021
- Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in TimeShaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu et al.CVPR 2021
- HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised LearningXiaozheng Zheng, Chao Wen, Zhou Xue, Pengfei Ren et al.ICCV 2023 · 17 citations
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
- Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ACM MM 2025 · 1 citation
