Lune

ACM MM2025Top-tier venue

Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose Estimation

Pengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi, Jing Wang, Jianxin Liao

2025Year
1Citations
3Top-tier citations

Abstract

Self-supervised 3D hand pose estimation methods can leverage labeled synthetic data along with unlabeled real-world data for model training, thereby alleviating the reliance on large-scale annotated datasets. Multi-view information fusion is a key factor in the success of these methods. Rule-based fixed fusion methods are simple, efficient, and generalizable, but they neglect the rich visual information in each view. Neural network-based learnable fusion methods can effectively model both intra- and inter-view semantic context, but they tend to overfit to the domain-specific feature of synthetic data and susceptible to interference of domain gaps. In this paper, we decompose multi-view fusion into two components: a learnable confidence estimation stage and a fixed confidence fusion stage. This design not only enables effective use of multi-view semantic cues but also ensures strong cross-domain generalization. To achieve accurate and robust confidence estimation, our method jointly exploits both multi-view pose consistency and pose-to-data consistency. Experiments on three public datasets demonstrate that our approach significantly outperforms existing state-of-the-art self-supervised 3D hand pose estimation methods.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers3

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines