HanDiffuser: Text-to-Image Generation with Realistic Hand Appearances
Supreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta, Saayan Mitra, Minh Hoai
摘要
Text-to-image generative models can generate highquality humans, but realism is lost when generating hands. Common artifacts include irregular hand poses, shapes, incorrect numbers of fingers, and physically implausible finger orientations. To generate images with realistic hands, we propose a novel diffusion-based architecture called HanDiffuser that achieves realism by injecting hand embeddings in the generative process. HanDiffuser consists of two components: a Text-to-Hand-Params diffusion model to generate SMPL-Body and MANO-Hand parameters from input text prompts, and a Text-Guided Hand-Params-to-Image diffusion model to synthesize images by conditioning on the prompts and hand parameters generated by the previous component. We incorporate multiple aspects of hand representation, including 3D shapes and joint-level finger positions, orientations and articulations, for robust learning and reliable performance during inference. We conduct extensive quantitative and qualitative experiments and perform user studies to demonstrate the efficacy of our method in generating images with high-quality hands. Project page: https://supreethn.github.io/ research/handiffuser/index.html * Work started when Supreeth was an intern at Adobe Research gers can bend to various degrees relatively independently. Hands can also occur in various shapes, sizes, and orientations and can be occluded with other human body parts. Further, hands often interact with objects and can have a wide range of grasps depending on the object's size, shape, and affordance. Therefore, capturing such a vast range of articulations and interactions directly from text inputs remains challenging. Despite having billions of parameters and several millions of trainable images, T2I models struggle to generate realistic hands.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Understanding Hallucinations in Diffusion Models through Mode InterpolationSumukh K. Aithal, Pratyush Maini, Zachary C. Lipton, J. Zico KolterNeurIPS 2024 · 被引用 121 次
- HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional InpaintingWenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang 等ACM MM 2024 · 被引用 20 次
- MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank ExpertsJie Zhu, Yixiong Chen, Mingyu Ding, Ping Luo 等NeurIPS 2024 · 被引用 17 次
- Uncovering Conceptual Blindspots in Generative Image Models Using Sparse AutoencodersMatyas Bohacek, Thomas Fel, Maneesh Agrawala, Ekdeep Singh LubanaICLR 2026 · 被引用 7 次
- X-Dancer: Expressive Music to Human Dance Video GenerationZeyuan Chen, Hongyi Xu, Guoxian Song, You Xie 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Hand1000: Generating Realistic Hands from Text with Only 1, 000 ImagesHaozhuo Zhang, Bin Zhu, Yu Cao, Yanbin HaoAAAI 2025 · 被引用 11 次
- HOIDiffusion: Generating Realistic 3D Hand-Object Interaction DataMengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu 等CVPR 2024 · 被引用 18 次
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng 等CVPR 2026 · 被引用 6 次
- HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human GenerationXin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang 等CVPR 2024
- RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated ImagesBenzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang 等AAAI 2025 · 被引用 10 次
