RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images
Benzhi Wang, Jingkai Zhou, Jingqi Bai, Yang Yang, Weihua Chen, Fan Wang, Zhen Lei
Abstract
In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intricate structural complexity. To address this issue, we propose a novel post-processing solution named RealisHuman. The RealisHuman framework operates in two stages. First, it generates realistic human parts, such as hands or faces, using the original malformed parts as references, ensuring consistent details with the original image. Second, it seamlessly integrates the rectified human parts back into their corresponding positions by repainting the surrounding areas to ensure smooth and realistic blending. The RealisHuman framework significantly enhances the realism of human generation, as demonstrated by notable improvements in both qualitative and quantitative metrics. Code is available at https://github.com/Wangbenzhi/RealisHuman .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c15ee224-00f2-4ebb-8c5f-00ebf5459ca3Cited by top-tier papers7
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li et al.CVPR 2026 · 3 citations
- SGMHand: Structure-Guided Modulation for Structure-Aware Hand InpaintingChuancheng Shi, Shiming Guo, Ke Shui, Yixiang Chen et al.AAAI 2026 · 1 citation
- FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image GenerationKefan Chen, Chaerin Min, Linguang Zhang, Shreyas Hampali et al.CVPR 2025
- Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentYizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim et al.ICLR 2025
- ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable GraspingYouxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu et al.CVPR 2025
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional InpaintingWenquan Lu, Yufei Xu, Jing Zhang, Chaoyue Wang et al.ACM MM 2024 · 20 citations
- HanDiffuser: Text-to-Image Generation with Realistic Hand AppearancesSupreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta et al.CVPR 2024 · 17 citations
- HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human GenerationXin Huang, Ruizhi Shao, Qi Zhang, Hongwen Zhang et al.CVPR 2024
- DiffHuman: Probabilistic Photorealistic 3D Reconstruction of HumansAkash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona et al.CVPR 2024 · 10 citations
- RHanDS: Refining Malformed Hands for Generated Images with Decoupled Structure and Style GuidanceChengrui Wang, Pengfei Liu, Min Zhou, Ming Zeng et al.AAAI 2025 · 12 citations
