FLASHand: Feed-forward reLightable and Animatable Single-view Hand Reconstruction
Ling-Xiao Zhang, Lin Gao, Wei-Hong He, Yu-Xuan Yang, Yunbing Xing, Yu-Kun Lai, Yiqiang Chen
摘要
Reconstructing high-fidelity, animatable, and relightable 3D hand avatars from a single RGB image is a challenging yet critical task for immersive VR/AR applications. State-of-the-art hand reconstruction methods achieve impressive reconstruction and relighting results, but they mostly leverage dense observations, such as multi-view images or monocular video sequences and rely on per-scene optimization. Moreover, it is difficult for these methods to generate plausible appearance in occluded regions. In contrast, existing single-view hand reconstruction methods typically struggle to disentangle global illumination, resulting in textures with baked-in shading artifacts. To address these challenges, we propose FLASHand, the first feed-forward relightable and animatable 3D hand reconstruction model from a single RGB image. Our method bridges the gap between efficiency and fidelity, enabling instant creation of personalized hand avatars with disentangled appearance that can be rendered under novel lighting and poses. To ensure plausible geometry generation and achieve high-fidelity animation and relighting, we introduce a mesh-based disentangled 2D Gaussian splatting hand representation. We leverage the NIMBLE parametric prior to define a canonical attribute space, and define geometry and appearance attributes explicitly. We then propose the Feed-forward Hand Attributes Generator (FHAG) to predict these attributes from a single image. By leveraging a cross attention module to fuse canonical geometric information with input image features, FHAG effectively lifts pixel-level visual cues into the 3D hand canonical space, directly generating spatially-aligned geometry and appearance attributes. This allows us to bypass iterative optimization and instantly reconstruct a high-fidelity hand avatar from a single RGB image. Extensive experiments on both synthetic and in-the-wild datasets demonstrate that FLASHand achieves state-of-the-art performance in novel view synthesis and supports real-time animation and relighting. Code and data are available at https://github.com/IGLICT/FLASHand.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Relightable and Dynamic Gaussian Avatar Reconstruction from Monocular VideoSeonghwa Choi, Moonkyeong Choi, Mingyu Jang, Jaekyung Kim 等ACM MM 2025 · 被引用 1 次
- Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand AvatarsXuan Huang, Hanhui Li, Wanquan Liu, Xiaodan Liang 等NeurIPS 2024 · 被引用 6 次
- Hand Avatar: Free-Pose Hand Animation and Rendering from Monocular VideoXingyu Chen, Baoyuan Wang, Heung-Yeung ShumCVPR 2023
- RelightAnyone: A Generalized Relightable 3D Gaussian Head ModelYingyan Xu, Pramod Rao, Sebastian Weiss, Gaspard Zoss 等CVPR 2026
- URHand: Universal Relightable HandsZhaoxi Chen, Gyeongsik Moon, Kaiwen Guo, Chen Cao 等CVPR 2024
