FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
Kefan Chen, Chaerin Min, Linguang Zhang, Shreyas Hampali, Cem Keskin, Srinath Sridhar
摘要
Figure 1 . We present FoundHand, a domain-specific image generation model that can synthesize realistic single and dual hand images. FoundHand is trained on our large-scale FoundHand-10M dataset which contains automatically extracted 2D keypoints and segmentation mask annotations (top left). FoundHand is formulated as a 2D pose-conditioned image-to-image diffusion model that enables precise hand pose and camera viewpoint control (top right). Optionally, we can condition the generation with a reference image to preserve its style (top right). Our model demonstrates robust in-the-wild generalization across hand-centric applications and has core capabilities such as gesture transfer, domain transfer, and novel view synthesis (middle row). This endows FoundHand with zero-shot applications to fix malformed hand images and synthesize coherent hand and hand-object videos, without explicitly giving object cues (bottom row).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang 等CVPR 2026 · 被引用 5 次
- EclipseTouch: Touch Segmentation on Ad Hoc Surfaces using Worn Infrared Shadow CastingVimal Mollyn, Nathan Devrio, Chris HarrisonUIST 2025 · 被引用 3 次
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li 等CVPR 2026 · 被引用 3 次
- A Temporal and Content Co-Awareness Latent Diffusion for Controllable Hand Image GenerationShuang Hao, Pengfei Ren, Haifeng Sun, Pan Ting 等CVPR 2026
- SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural AlignmentZhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei 等ICLR 2026
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu 等CVPR 2024 · 被引用 12 次
- UniHand: A Unified Model for Diverse Controlled 4D Hand Motion ModelingZhihao Sun, Tong Wu, Ruirui Tu, Daoguo Dong 等ICLR 2026 · 被引用 2 次
- HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional InpaintingWenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang 等ACM MM 2024 · 被引用 20 次
- HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point CloudWencan Cheng, Hao Tang, Luc Van Gool, Jong Hwan KoCVPR 2024
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
