FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
Kefan Chen, Chaerin Min, Linguang Zhang, Shreyas Hampali, Cem Keskin, Srinath Sridhar
Abstract
Figure 1 . We present FoundHand, a domain-specific image generation model that can synthesize realistic single and dual hand images. FoundHand is trained on our large-scale FoundHand-10M dataset which contains automatically extracted 2D keypoints and segmentation mask annotations (top left). FoundHand is formulated as a 2D pose-conditioned image-to-image diffusion model that enables precise hand pose and camera viewpoint control (top right). Optionally, we can condition the generation with a reference image to preserve its style (top right). Our model demonstrates robust in-the-wild generalization across hand-centric applications and has core capabilities such as gesture transfer, domain transfer, and novel view synthesis (middle row). This endows FoundHand with zero-shot applications to fix malformed hand images and synthesize coherent hand and hand-object videos, without explicitly giving object cues (bottom row).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07eb4b68-ada1-496a-8774-71b21f2334eaCited by top-tier papers5
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang et al.CVPR 2026 · 5 citations
- EclipseTouch: Touch Segmentation on Ad Hoc Surfaces using Worn Infrared Shadow CastingVimal Mollyn, Nathan Devrio, Chris HarrisonUIST 2025 · 3 citations
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li et al.CVPR 2026 · 3 citations
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting et al.CVPR 2026
- SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural AlignmentZhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei et al.ICLR 2026
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu et al.CVPR 2024 · 12 citations
- UniHand: A Unified Model for Diverse Controlled 4D Hand Motion ModelingZhihao Sun, Tong Wu, Ruirui Tu, Daoguo Dong et al.ICLR 2026 · 2 citations
- HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional InpaintingWenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang et al.ACM MM 2024 · 20 citations
- HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point CloudWencan Cheng, Hao Tang, Luc Van Gool, Jong Hwan KoCVPR 2024
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
