Hand1000: Generating Realistic Hands from Text with Only 1, 000 Images
Haozhuo Zhang, Bin Zhu, Yu Cao, Yanbin Hao
摘要
Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations of human hands. The resulting images frequently exhibit issues such as incorrect numbers of fingers, unnatural twisting or interlacing of fingers, or blurred and indistinct hands. These issues stem from the inherent complexity of hand structures and the difficulty in aligning textual descriptions with precise visual depictions of hands. To address these challenges, we propose a novel approach named Hand1000 that enables the generation of realistic hand images with target gesture using only 1,000 training samples. The training of Hand1000 is divided into three stages with the first stage aiming to enhance the model's understanding of hand anatomy by using a pre-trained hand gesture recognition model to extract gesture representation. The second stage further optimizes text embedding by incorporating the extracted hand gesture representation, to improve alignment between the textual descriptions and the generated hand images. The third stage utilizes the optimized embedding to fine-tune the Stable Diffusion model to generate realistic hand images. In addition, we construct the first publicly available dataset specifically designed for text-to-hand image generation. Based on the existing hand gesture recognition dataset, we adopt advanced image captioning models and LLaMA3 to generate high-quality textual descriptions enriched with detailed gesture information. Extensive experiments demonstrate that Hand1000 significantly outperforms existing models in producing anatomically correct hand images while faithfully representing other details in the text, such as faces, clothing and colors. Additional details and resources are available on our project page: https://haozhuo-zhang.github.io/Hand1000-project-page/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang 等CVPR 2026 · 被引用 5 次
- RAGG: Retrieval-Augmented Grasp Generation ModelZhenhua Tang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo 等AAAI 2025 · 被引用 2 次
- ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable GraspingYouxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu 等CVPR 2025
- LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech RecognitionFeng Xue, Baochao Zhu, Wei Jia, Shujie Li 等AAAI 2026
- SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural AlignmentZhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei 等ICLR 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- HanDiffuser: Text-to-Image Generation with Realistic Hand AppearancesSupreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta 等CVPR 2024 · 被引用 17 次
- Text-Driven 3D Hand Motion Generation from Sign Language DataLéore Bensabath, Mathis Petrovich, Gül VarolCVPR 2026 · 被引用 5 次
- FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image GenerationKefan Chen, Chaerin Min, Linguang Zhang, Shreyas Hampali 等CVPR 2025
- MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank ExpertsJie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo 等NeurIPS 2024 · 被引用 17 次
- Handy: Towards a High Fidelity 3D Hand Shape and Appearance ModelRolandos Alexandros Potamias, Stylianos Ploumpis, Stylianos Moschoglou, Vasileios Triantafyllou 等CVPR 2023
