FLASHand: Feed-forward reLightable and Animatable Single-view Hand Reconstruction
Ling-Xiao Zhang, Lin Gao, Wei-Hong He, Yu-Xuan Yang, Yunbing Xing, Yu-Kun Lai, Yiqiang Chen
Abstract
Reconstructing high-fidelity, animatable, and relightable 3D hand avatars from a single RGB image is a challenging yet critical task for immersive VR/AR applications. State-of-the-art hand reconstruction methods achieve impressive reconstruction and relighting results, but they mostly leverage dense observations, such as multi-view images or monocular video sequences and rely on per-scene optimization. Moreover, it is difficult for these methods to generate plausible appearance in occluded regions. In contrast, existing single-view hand reconstruction methods typically struggle to disentangle global illumination, resulting in textures with baked-in shading artifacts. To address these challenges, we propose FLASHand, the first feed-forward relightable and animatable 3D hand reconstruction model from a single RGB image. Our method bridges the gap between efficiency and fidelity, enabling instant creation of personalized hand avatars with disentangled appearance that can be rendered under novel lighting and poses. To ensure plausible geometry generation and achieve high-fidelity animation and relighting, we introduce a mesh-based disentangled 2D Gaussian splatting hand representation. We leverage the NIMBLE parametric prior to define a canonical attribute space, and define geometry and appearance attributes explicitly. We then propose the Feed-forward Hand Attributes Generator (FHAG) to predict these attributes from a single image. By leveraging a cross attention module to fuse canonical geometric information with input image features, FHAG effectively lifts pixel-level visual cues into the 3D hand canonical space, directly generating spatially-aligned geometry and appearance attributes. This allows us to bypass iterative optimization and instantly reconstruct a high-fidelity hand avatar from a single RGB image. Extensive experiments on both synthetic and in-the-wild datasets demonstrate that FLASHand achieves state-of-the-art performance in novel view synthesis and supports real-time animation and relighting. Code and data are available at https://github.com/IGLICT/FLASHand.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e7be40ff-bc23-4951-99c9-d263cbedd7caRelated papers
- Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular VideoSeonghwa Choi, Moonkyeong Choi, Mingyu Jang, Jaekyung Kim et al.ACM MM 2025 · 1 citation
- Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand AvatarsXuan Huang, Hanhui Li, Wanquan Liu, Xiaodan Liang et al.NeurIPS 2024 · 6 citations
- Hand Avatar: Free-Pose Hand Animation and Rendering from Monocular VideoXingyu Chen, Baoyuan Wang, Heung-Yeung ShumCVPR 2023
- RelightAnyone: A Generalized Relightable 3D Gaussian Head ModelYingyan Xu, Pramod Rao, Sebastian Weiss, Gaspard Zoss et al.CVPR 2026
- URHand: Universal Relightable HandsZhaoxi Chen, Gyeongsik Moon, Kaiwen Guo, Chen Cao et al.CVPR 2024
