Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba
Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente Carrasco, Fernando De la Torre
摘要
3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, primarily due to inefficiently modeling spatial relations between joints. To address this problem, we propose a novel graph-guided Mamba framework, named Hamba, which bridges graph learning and state space modeling. Our core idea is to reformulate Mamba's scanning into graph-guided bidirectional scanning for 3D reconstruction using a few effective tokens. This enables us to efficiently learn the spatial relationships between joints for improving reconstruction performance. Specifically, we design a Graph-guided State Space (GSS) block that learns the graph-structured relations and spatial sequences of joints and uses 88.5% fewer tokens than attention-based methods. Additionally, we integrate the state space features and the global features using a fusion module. By utilizing the GSS block and the fusion module, Hamba effectively leverages the graph-guided state space features and jointly considers global and local features to improve performance. Experiments on several benchmarks and in-the-wild tests demonstrate that Hamba significantly outperforms existing SOTAs, achieving the PA-MPVPE of 5.3mm and F@15mm of 0.992 on FreiHAND. At the time of this paper's acceptance, Hamba holds the top position, Rank 1 in two Competition Leaderboards on 3D hand reconstruction. Project Website: https://humansensinglab.github.io/Hamba/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- Long-LRM: Long-Sequence Large Reconstruction Model for Wide-Coverage Gaussian SplatsZiwen Chen, Hao Tan, Kai Zhang, Sai Bi 等ICCV 2025 · 被引用 14 次
- MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion GenerationBohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing LuNeurIPS 2025 · 被引用 14 次
- GUAVA: Generalizable Upper Body 3D Gaussian AvatarDongbin Zhang, Yunfei Liu, Lijian Lin, Ye Zhu 等ICCV 2025 · 被引用 9 次
- HoliGS: Holistic Gaussian Splatting for Embodied View SynthesisXiaoyuan Wang, Yizhou Zhao, Botao Ye, Xiaojun Shan 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper45
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- PoseMamba: Monocular 3D Human Pose Estimation with Bidirectional Global-Local Spatio-Temporal State Space ModelYunlong Huang, Junshuo Liu, Ke Xian, Robert Caiming QiuAAAI 2025 · 被引用 15 次
- Disentangled Hypergraph-Guided Mamba Scanning for Fine-Grained Visual RecognitionZhongwei Xiong, Hao Wang, Xiaoyan Yu, Lingling Li 等AAAI 2026
- Samba: A Unified Mamba-based Framework for General Salient Object DetectionJiahao He, Keren Fu, Xiaohong Liu, Qijun ZhaoCVPR 2025
- Pose Magic: Efficient and Temporally Consistent Human Pose Estimation with a Hybrid Mamba-GCN NetworkXinyi Zhang, Qiqi Bao, Qinpeng Cui, Wenming Yang 等AAAI 2025 · 被引用 18 次
- HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose EstimationWencan Cheng, Gim Hee LeeAAAI 2026
