JUMP-Hand: Learning Joint-wise Uncertainty to Gate Mixture of View Experts for Multi-View 3D Hand Reconstruction
Haohong Kuang, Yang Xiao, Changlong Jiang, Jinghong Zheng, Hang Xu, Ran Wang, Zhiguo Cao, Joey Tianyi Zhou
Abstract
We propose JUMP-Hand, a novel multi-view 3D hand reconstruction method that explicitly models probabilistic joint-wise uncertainty as a gating mechanism for multi-view fusion. Existing approaches usually rely on naive pooling or implicit attention, overlooking that each hand joint exhibits varying visibility and reliability across views. Such indiscriminate aggregation of noisy or unreliable information inevitably degrades overall performance. For instance, a joint severely occluded in one view might be clearly visible in another. To address this, JUMP-Hand draws inspiration from the Mixture of Experts (MoE) paradigm, treating each 2D view as a specialized expert. The key idea is that the reliability of each view expert is quantified through joint-wise uncertainty modeling, serving as an explicit gating signal to route experts' partial yet complementary clues for each joint in a coarse-to-fine reconstruction paradigm. In this design, uncertainty not only guides the uncertainty-aware triangulation for reliable 3D hand initialization during the coarse stage, but also acts as a gating signal during the refinement stage to adaptively aggregate multi-scale features from different view experts on a joint-wise basis, enabling robust 3D hand reconstruction. Extensive experiments on DexYCB-MV, HO3D-MV, and OakInk-MV demonstrate that our method consistently achieves state-of-the-art results, validating the effectiveness of the proposed method with joint-wise uncertainty gating for reliable 3D hand reconstruction. Code is available at https:// github.com/ HaohongKuang/ JUMP-Hand.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91db9838-6e6d-4242-ab10-46c75747ee3aBuilds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth ImageFu Xiong, Boshen Zhang, Yang Xiao, Zhiguo Cao et al.ICCV 2019 · 178 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
Related papers
- H2ONet: Hand-Occlusion-and-Orientation-Aware Network for Real-Time 3D Hand Mesh ReconstructionHao Xu, Tianyu Wang, Xiao Tang, Chi-Wing FuCVPR 2023
- MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic GuidanceHang Xu, Yang Xiao, Changlong Jiang, Haohong Kuang et al.AAAI 2026
- MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsJingnan Gao, Zhe Wang, Xianze Fang, Xingyu Ren et al.CVPR 2026 · 19 citations
- GeoMoE: Divide-and-Conquer Motion Field Modeling with Mixture-of-Experts for Two-View GeometryJiajun Le, Jiayi MaAAAI 2026
- A Probabilistic Attention Model with Occlusion-aware Texture Regression for 3D Hand Reconstruction from a Single RGB ImageZheheng Jiang, Hossein Rahmani, Sue Black, Bryan M. WilliamsCVPR 2023
