Mining Multi-View Information: A Strong Self-Supervised Framework for Depth-based 3D Hand Pose and Mesh Estimation
Pengfei Ren, Haifeng Sun, Jiachang Hao, Jingyu Wang, Qi Qi, Jianxin Liao
Abstract
In this work, we study the cross-view information fusion problem in the task of self-supervised 3D hand pose estimation from the depth image. Previous methods usually adopt a hand-crafted rule to generate pseudo labels from multi-view estimations in order to supervise the network training in each view. However, these methods ignore the rich semantic information in each view and ignore the complex dependencies between different regions of different views. To solve these problems, we propose a cross-view fusion network to fully exploit and adaptively aggregate multi-view information. We encode diverse semantic information in each view into multiple compact nodes. Then, we introduce the graph convolution to model the complex dependencies between nodes and perform cross-view information interaction. Based on the cross-view fusion network, we propose a strong self-supervised framework for 3D hand pose and hand mesh estimation. Furthermore, we propose a pseudo multi-view training strategy to extend our framework to a more general scenario in which only single-view training data is used. Results on NYU dataset demonstrate that our method outperforms the previous self-supervised methods by 17.5% and 30.3% in multi-view and single-view scenarios. Meanwhile, our framework achieves comparable re-sults to several strongly supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 290538b3-f4a2-4da3-8e44-3e4c798f472fCited by top-tier papers10
- Two Heads Are Better than One: Image-Point Cloud Network for Depth-Based 3D Hand Pose EstimationPengfei Ren, Yuchen Chen, Jiachang Hao, Haifeng Sun et al.AAAI 2023 · 28 citations
- Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB ImagePengfei Ren, Chao Wen, Xiaozheng Zheng, Zhou Xue et al.ICCV 2023 · 15 citations
- Touchscreen-based Hand Tracking for Remote Whiteboard InteractionXinshuang Liu, Yizhong Zhang, Xin TongUIST 2024 · 8 citations
- Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action RecognitionHaochen Chang, Pengfei Ren, Haoyang Zhang, Liang Xie et al.ICCV 2025 · 8 citations
- Dynamic Support Information Mining for Category-Agnostic Pose EstimationPengfei Ren, Yuanyuan Gao, Haifeng Sun, Qi Qi et al.CVPR 2024 · 3 citations
Builds on21
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang et al.ICCV 2019 · 242 citations
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageFu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao et al.ICCV 2019 · 178 citations
- Perturbed Self-Distillation: Weakly Supervised Large-Scale Point Cloud Semantic SegmentationYachao Zhang, Yanyun Qu, Yuan Xie, Zonghao Li et al.ICCV 2021 · 138 citations
Related papers
- Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ACM MM 2025 · 1 citation
- HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised LearningXiaozheng Zheng, Chao Wen, Zhou Xue, Pengfei Ren et al.ICCV 2023 · 17 citations
- HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth MapJameel Malik, Ibrahim Abdelaziz, Ahmed Elhayek, Soshi Shimada et al.CVPR 2020
- Efficient Virtual View Selection for 3D Hand Pose EstimationJian Cheng, Yanguang Wan, Dexin Zuo, Cuixia Ma et al.AAAI 2022 · 30 citations
- Deep Semantic Graph Transformer for Multi-View 3D Human Pose EstimationLijun Zhang, Kangkang Zhou, Feng Lu, Xiang-Dong Zhou et al.AAAI 2024 · 14 citations
